
Gergely Orosz
@GergelyOrosz • 352,080 subscribers
Writing @Pragmatic_Eng, the #1 software engineering newsletter on Substack. Author of @EngGuidebook. Formerly Uber & Skype.
Videos

If you’ve ever opened Chrome DevTools, or optimized a page for Core Web Vitals, you’ve used software built by Addy Osmani. Timestamps: 00:00 Intro 02:50 Addy’s current workflow 05:11 Addy’s path into tech 15:04 Addy’s work on jQuery 16:44 TodoMVC 21:44 Getting hired at Google and working on Chrome 27:17 Building dev tools 40:15 Core Web Vitals 45:42 Google’s engineering culture 51:03 Addy’s career trajectory at Google 57:55 The director role at Google 1:01:40 Cognitive debt and cognitive surrender 1:03:03 Working with agents 1:05:52 Loop engineering 1:12:55 The changing role of the software engineer 1:18:15 How Addy uses AI in writing 1:27:40 What’s next for Addy 1:28:47 Career advice Brought to you by: • Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. Teams like Jane Street, and the etcd community use Antithesis to ship better code, faster. • Sentry – application monitoring software considered “not bad” by millions of developers. • Google Cloud Run – run untrusted agent code without the security anxiety. Cloud Run sandboxes deliver hyper-isolated, ephemeral execution environments that spin up in milliseconds. Check them out: Here's Addy's advice on where he believes engineers should invest efforts, in the coming years, in his words: “What we are very likely to see happen next with engineering careers (as well as product and other roles) is the unbundling of them, so that an engineer also has product sense, while a product person also has engineering sense, or UX sense. You should think about the non-engineering things if you don’t [usually] have the time to think about product or technical evangelism, or go-to-market approaches, or any other parts of how businesses are successful. If you can show employers that you are not just a builder, but someone that can help them as roles start to become a little bit fuzzier, then I think that you can be successful in these times. Don’t be just an engineer.”
Gergely Orosz398,385 views • 15 days ago

Few people care more about software performance than Casey Muratori. Give him a few minutes of your attention with this episode, and he'll convince you to learn to read Assembly (no, really, I finally started to read it, it's really not that scary, esp with an AI that can help explain the sequences). Timestamps: 00:00 Intro 05:17 Games at Microsoft 12:52 Building games 16:00 Why performance matters 27:12 Why you should learn to read assembly 30:36 Designing for optimization 42:51 How to get better at writing performant software 49:04 Understanding how the CPU works 55:53 Building games then and now 1:05:56 How game engines changed building games 1:10:48 Why new games compete with old games 1:13:25 GTA 6: why is it taking so long? 1:16:59 Casey's critique of clean code 1:21:48 Casey's take on TDD 1:24:30 What is good code? 1:27:32 What makes a good software engineer? 1:33:56 Why Casey doesn't code with AI 1:39:01 AI's impact on the game industry 1:44:43 AI and burnout 1:50:21 Why you should read papers Brought to you by: • Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages • Sentry – application monitoring software considered “not bad” by millions of developers • turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable Three interesting things we talked about: 1. Is performance starting to matter to businesses? Enterprise software buyers care mainly about cost, compliance, and capabilities – but not performance. Even so, there are some products gaining major popularity and market share due to their performance, such as File Pilot (next-gen file explorer) and the Blick video editor. Is the tide turning? 2. Profiler-driven performance optimization is the wrong way to optimize The standard way of optimizing is to profile the application, tweak hotspots, then check if the stats have improved. But this only finds a local minimum; Casey says every engineer he’s worked with who was a great “optimizer” began by establishing what the hardware could theoretically do, and then did not stop until they’d closed the gap to that performance level. 3. Take a grain of salt with conventional wisdom that premature optimization is the “root of all evil” Many devs use it as an excuse to delay performance optimization, but Casey says that not optimizing in time could mean that only performance hotspots can be fixed later, and not the architectural issues that create poor performance. Architect your system to be performant, or you’ll have trouble solving problems without a rewrite!
Gergely Orosz95,079 views • 8 days ago

On Bending Spoons: I still think about when they bought Evernote in 2023, it had user data stored on 750 manually provisioned (!!) VMs on GCP, sharded, with ongoing performance + reliability issues, running a Java 11 monolith (!!!) Then they went ahead and fixed it in ~6 months
Gergely Orosz186,804 views • 29 days ago

What does it mean for software engineering when we no longer write the code? Here's the take from Boris Cherny (Boris Cherny), the creator of Claude Code. Timestamps: 00:00 Intro 11:15 Lessons from Meta 19:46 Joining Anthropic 23:08 The origins of Claude Code 32:55 Boris's Claude Code workflow 36:27 Parallel agents 40:25 Code reviews 47:18 Claude Code's architecture 52:38 Permissions and sandboxing 55:05 Engineering culture at Anthropic 1:05:15 Claude Cowork 1:12:48 Observability and privacy 1:14:45 Agent swarms 1:21:16 LLMs and the printing press analogy 1:30:16 Standout engineer archetypes 1:32:12 What skills still matter for engineers 1:35:24 Book recommendations Brought to you by: • Statsig — The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for automated code review. Proactively find and fix issues in real-time with the SonarQube MCP Server: • WorkOS – Everything you need to make your app enterprise ready. Three interesting things from this conversation: 1. Boris automated himself out of code review well before AI. Boris was one of the most prolific code reviewers at Meta company. And he worked hard to minimize time spent on code review. His system::every time he left the same kind of review comment, he logged it in a spreadsheet. Once a pattern hit 3-4 occurrences, he’d write a lint rule to automate it away! 2. PRDs are dead on the Claude Code team: prototypes replaced them. Instead of writing Product Requirement Documents (specs), they build hundreds of working prototypes before shipping a feature. Boris: “There’s just no way we could have shipped this if we started with static mocks and Figma or if we started with a PRD.” 3. This is the year of the generalist (and maybe the year of those with ADHD) Boris’s work has shifted from deep-focus single-threaded coding to managing multiple parallel agents and context-switching rapidly. As Boris put it: “It’s not so much about deep work, it’s about how good I am at context switching and jumping across multiple different contexts very quickly.”
Gergely Orosz490,747 views • 6 months ago

Why is the creator of OpenCode pretty skeptical about AI productivity gains, and the hype around AI? A very conversation dax (and lots of truth bombs:) Timestamps: 00:00 Intro 07:03 Dax’s path into tech 09:04 Early startup experience 13:16 Getting involved with open source 16:13 OpenCode 23:17 Anthropic banning OpenCode 30:34 From terminal to GUI 32:34 OpenCode’s business model 36:33 Why inference is profitable 39:11 GPU bottlenecks 40:54 AI hype 45:50 AI spending 48:47 Dax’s memo 55:41 Dax’s skepticism of predictions 58:58 Engineering culture at OpenCode 1:02:38 How building works at OpenCode 1:05:36 Taste and quality 1:11:32 Dax’s work setup 1:12:35 The role of engineers and EMs 1:15:50 Advice for engineers 1:18:12 Book recommendation Brought to you by: • Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages • WorkOS – everything you need to make your app enterprise ready • turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable Three interesting thoughts from Dax: 1. No AI-native coding agent company is “winning” by being better with AI. Dax says that none of OpenCode’s competitors are crushing them, and that nobody is using AI so well that others cannot compete. 2. Most software engineers profit from AI as time gained, not increased output — unless you change incentives! Dax says the natural way for software engineers to “cash out” their AI tooling gains is with time savings, by doing the same work as before, but faster. Until compensation and motivation structures change, most teams should expect output to stay flat while engineers go home earlier. There’s nothing wrong with this, but AI vendors sell a different outcome to CFOs: increased output. 3. AI code generation mutes the “guilt” of doing the wrong thing, but this builds up tech debt. Pre-AI, writing a hack felt bad, the second time it felt really bad, and by the third time you’d often just refactor in order to fix up the code. Now, the agent hides the hack, which skews devs’ judgment and results in less tech debt being cleaned up.
Gergely Orosz231,584 views • 3 months ago

Just six months ago, DHH (creator of Ruby on Rails and Omarchy) said how he doesn’t really use AI tools to write code, because they are not good enough. Things have changed, a lot. Timestamps: 00:00 Intro 02:11 Omarchy and Ruby on Rails 08:25 37signals overview 10:12 Launching HEY 18:38 Building HEY 22:47 Designers at 37signals 28:08 The craft of design 31:52 Why DHH now embraces AI workflows 39:45 The AI inflection point 44:23 DHH’s agent-first workflow 55:09 AI’s impact on junior developers 1:03:08 Developer experience with AI 1:16:43 What does AI mean for developers? 1:23:33 37signals teams and hiring 1:38:20 Work-life balance with AI 1:41:41 Why DHH keeps building 1:45:24 Closing Brought to you by: • Statsig – The unified platform for flags, analytics, experiments, and more. Stop switching between different tools, and have them all in one place. • WorkOS – Everything you need to make your app enterprise ready. WorkOS gives you APIs to ship enterprise features in days. Check out • Sonar – The makers of SonarQube, the industry standard for automated code review. See how SonarQube Advanced Security is empowering the Agent Centric Development Cycle (AC/DC) with new capabilities. Three interesting observations from this conversation: #1 DHH's philosophy on AI has not changed, but the available tools very much have. Autocomplete-style coding assistants were genuinely annoying for experienced developers six months ago. Things changed with the shift from tab-completion to agent harnesses, plus the emergence of powerful models like Opus 4.5 – when agents started producing code which DHH does want to merge with little to no alteration. #2 Beautiful code and products aren’t matters of vanity; they’re signals of correctness. Dipping into philosophy, DHH says: “When something is beautiful, it’s likely to be correct.” He argues that Steve Jobs wanted the inside of a computer to be beautiful because people who care about circuit board layout are also those who sweat on the details of the UI. #3 DHH’s development workflow, today: He runs tmux to have two models running, and neovim in the center. Specifics: - One fast LLM running (typically Gemini 2.5) in one split terminal - A slow but more powerful model in another terminal (usually Opus) - NeoVim for reviewing diffs via Lazygit
Gergely Orosz329,135 views • 4 months ago

OpenClaw - the agentic software spreading like wildfire - was built on top of Pi, a minimalist, self-modifying agent. I sat down with Pi's creator, Mario Zechner and longtime Pi user (+ the creator of Flask) Armin Ronacher ⇌ to talk Pi, and their (very grounded!) takes on building with AI. Timestamps: 00:00 Intro 07:30 How Mario, Armin, and Peter Steinberger met 15:15 How 30 dev teams use AI agents: learnings 21:50 The importance of judgment 24:26 Challenges when non-engineers write code 28:30 Downsides of over-automation 32:18 Pi 48:09 OpenClaw + Pi 50:54 “Clankers” 57:32 Open source and AI 1:00:22 Complexity as the enemy 1:02:50 Building an AI-native startup 1:11:52 “Slow the F down” 1:16:40 MCPs vs. CLI 1:25:03 Predictions and staying up to date • YouTube: • Spotify: • Apple: Brought to you by: • Statsig – The unified platform for flags, analytics, experiments, and more. • Sonar — The makers of SonarQube, the industry standard for code verification and automated code review. Try it out for yourself. • WorkOS – WorkOS gives you APIs to ship enterprise features – SSO, directory sync, RBAC, audit logs – in days, not months. Visit learn more. --- Three parts I found especially interesting in this discussion: 1. New trend: AI makes it harder for senior engineers to reject pointless complexity. Historically, senior engineers kept software complexity at bay simply by saying “no” a lot. But Armin observes that these days, more junior engineers and product managers deploy agent-scripted counterarguments when a senior colleague kicks an idea to the curb. This makes decision-making exhausting, and more bad ideas make it into production as a result. 2. It should be MUCH easier to build specialized tools for specific tasks. Different projects need different harness types because, as Mario points out, the same hammer is not ideal for every single construction job. As such, Pi is built with the goal of allowing the creation of specialized harnesses. It can modify itself so that a user can create the bespoke harness needed for any task. Mario believes it’s a preview of how self-modifiable software might look in the future. 3. Automation bias is one of the biggest risks of working with AI agents. Once devs confirm that an AI agent can produce acceptable code, they start to review its output less often, even though agents can – and do! – produce slop. Mario advises being far more sceptical with agents, and cautions that the quality of their output isn’t guaranteed, however well they performed previously.
Gergely Orosz172,936 views • 4 months ago

In 2025, it was rational to be skeptical about whether AI would change the future of software development. In 2026, it's not, anymore. With Charity Majors: Timestamps: 00:00 Intro 02:56 How Parse led to Honeycomb 06:00 The limits of individual productivity metrics 09:08 How Charity’s perspective on AI has evolved 13:50 Rewriting code vs. editing code 19:20 Production as a stage of development 22:14 Code reviews 26:56 Non-deterministic systems 31:11 Sensible uses of AI 37:41 The two AI camps 44:40 Why AI works so well for building software 49:42 DevOps 55:13 Modern observability 1:00:40 Handling context overload 1:01:56 What’s new in Observability Engineering’s 2nd edition 1:07:45 What effective leadership looks like 1:10:25 Engineering management: what is changing? 1:16:31 Junior engineers 1:18:01 AI fatigue 1:21:39 Book recommendations Brought to you by: • Antithesis — turbocharge testing of your systems by running your whole system under aggressive fault injection. Teams like Jane Street, and the etcd community rely on Antithesis. • Buildkite — the CI platform trusted by OpenAI, Anthropic, Cursor, Meta, Uber, NVIDIA, Airbnb and many more. Engineered to absorb whatever your coding agents throw at the build queue. • WorkOS — make your app and agents Enterprise Ready, with SSO, SCIM, RBAC, and more. 1. The question engineers need to answer: what would it take for you to be fully comfortable shipping code you have not read? Charity believes it is a “when” and not an “if” that professional software engineers will ship code they never looked at – and thus do not understand – to production. Engineering is building the systems that validate this code, and allow shipping with full confidence. 2. AI could have the software industry go through the “pets” to “cattle” change that compute infra went through in the 2010s. Up to now, writing software from scratch was far more expensive than editing existing software. But now, generating hundreds of variants of a function can be done faster than how long it would take you to hand-write it once. Charity believes that we might be at the beginning of the transition from “pets” to “cattle” that happened at the hardware infrastructure layer. Before the 2010s, configuring and repairing individual servers was commonly done. But with tools like Terraform and Kubernetes, individual servers having issues are no longer fixed up: they are re-created instead. Charity thinks the same might happen with code, sooner rather than later. When there’s an issue with the code, generate new code that solves it, and is verifyably correct. 3. Non-deterministic systems require more engineering discipline versus before. With code written by AI, we’re reducing the trust in the code (because we no longer wrote it), so we need to increase trust at the other part of the development process. Specifically, at validation: with things like tests, evals, and conformance testing.
Gergely Orosz30,495 views • 21 days ago

Knowing how LLM contexts work and how to work around context limitations – aka “context engineering” – is becoming so important. No better person to explain than dex Timestamps: 00:00 Intro 01:33 Dex’s path into tech 03:34 Early work in platform engineering 05:28 Replicated 11:24 Metalytics 12:36 12-factor agents 18:27 Context engineering 23:38 Harness engineering 26:11 Context overload 30:45 Loop engineering 44:34 Software factories before and after AI 50:33 Automation limits 55:18 Three options for automating 59:00 RPI framework 1:04:16 Intentional compaction 1:11:48 Token harder vs. token smarter 1:16:44 AI slop 1:19:15 HumanLayer 1:29:09 Book recommendation Brought to you by: • Antithesis — with Antithesis, you can use AI agents to work on critical systems without worrying about correctness. Teams like Jane Street, and the etcd community use Antithesis to ship better code, faster. • Buildkite — the CI orchestration platform built for reliable scale. Used by OpenAI, Anthropic, Cursor, Meta, Uber, Ramp, Nvidia, Airbnb and many more. • Sentry — application monitoring software built by developers, for developers. Check out their AI agent, Seer AI, and Sentry MCP. Three interesting learnings from this episode: 1. Lesson learned: Shipping unread code spells disaster within months. Dex experimented with having the model write the code and humans not reviewing anything in July 2025. Four months later, they shut things down and threw the whole system out. Production broke, and no matter how much the team prompted Opus 4.1, the model could not find the root cause. Once fixed, it took three weeks (!!) to re-onboard to a codebase no human had ever read 2. Context engineering 101: figure out where the “dumb zone” begins. As a rule of thumb, the less of the context window that is used, the better the outcomes are. This is because the attention mechanism is quadratic: the more that goes into the context window, the more compute is required to process it all. 3. “You’re completely right!” or “you’re right to push back on that” are phrases that mean it’s time to start a new session. These responses mean the LLM session is trajectory-poisoned, and you’re wasting time and tokens to continue. This is because models are autoregressive.
Gergely Orosz63,174 views • 1 month ago

Anders Hejlsberg (Anders Hejlsberg) is a living legend: he created Turbo Pascal, Delphi, C# and TypeScript (and today TypeScript is the most-used programming language, globally, as per GitHub.) Timestamps: 00:00 Intro 02:48 How Anders got into programming 05:40 Building his first compiler 07:44 Turbo Pascal 12:25 Delphi 14:53 Joining Microsoft 19:41 Building C# 29:11 Async/await 34:01 The rise of JavaScript 37:52 Building TypeScript 42:58 How the TypeScript compiler works 48:30 JavaScript’s strengths and weaknesses 52:18 How Anders uses AI 56:03 What language features work well with AI 1:02:49 How software craftsmanship is changing 1:07:49 Performance and efficiency 1:09:29 Anders’ tool stack 1:11:30 A 30-year career at Microsoft 1:13:40 Book recommendation Brought to you by: Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. WorkOS – Everything you need to make your app enterprise ready. turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable. Four things that stood out to me: 1. “10x better for 1/10th of the price” is a proven winner. This is what Turbo Pascal did: it sold for $49.95 when competing compilers cost $500, and it was faster and more interactive than competitors’ products. Conveniently, the low price tag also killed off piracy 2. C# might have not existed without a famous court case. Microsoft originally hired Anders to architect its Java tools (Visual J++), but the Sun versus Microsoft lawsuit (1997-2001) meant Microsoft could not build on top of Java, as the company that owned Java’s IP (Sun) sued MS for alleged unauthorized changes to the Java language. Microsoft realized it had to build a new language that combined VB’s productivity with C++’s power. This led to C# and .NET. 3. TypeScript exists because Anders refused to build Script# for the Outlook .com team. Microsoft’s Outlook .com team asked Anders’ C# team to productize “ScriptSharp,” a language to cross-compile C# to JavaScript. Anders and the C# team pushed back, suggesting that a better approach was to fix JavaScript. Anders felt strongly that to be attractive to the best-of-breed developers in the JavaScript ecosystem, you want people to write JavaScript, and not another language like C#. 4. Designing a programming language is a 10-year play. As Anders puts it: “Version one is great, but has all sorts of issues. You’ve got to do version two, but it’s not until version three that it really starts to be great. Then you’ve got to convince people to adopt it.”
Gergely Orosz129,558 views • 3 months ago

No project has gotten more traction in such a short time than by Peter Steinberger 🦞 But how is he building it? Watch or listen: • YouTube: • Spotify: • Apple: Brought to you by: • Statsig — The unified platform for flags, analytics, experiments, and more. Join us at The Pragmatic Summit I’m hosting with Statsig, on 11 February: • Sonar – The makers of SonarQube, the industry standard for automated code review. Join me online at the Sonar Summit, on 3rd March: • WorkOS – Everything you need to make your app enterprise ready. If you're in SF on 9 February, stop by at the WorkOS AI Night with The Pragmatic Engineer (free to register):
Gergely Orosz225,804 views • 7 months ago

There’s a popular theory that AI will finally make formal verification mainstream because mathematical proof of correctness will be needed when machines write most or all of the code. But will this happen? Hillel Wayne is one of the best people to answer. Timestamps: 00:00 Intro 04:32 The Crossover Project 11:37 What software engineering does better 15:30 What traditional engineering does better 18:17 Formal methods 29:32 TLA+: what it is and demo 36:58 TLA+ at Amazon 38:10 Ways distributed systems break 41:03 Formal methods and systems thinking 46:20 The value of learning math 50:23 What TLA+ is good for and isn’t 52:50 Alloy: a declarative language for software modeling 58:53 Other formal methods tools 1:01:24 Property-based testing 1:05:31 AI and the need for formal verification 1:12:29 Logic for programmers 1:14:35 Hillel’s 2025 prediction on AI’s impact 1:21:30 Book recommendation Brought to you by: • Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. • turbopuffer – a vector and full-text search engine built on object storage. It’s fast, cheap, and extremely scalable. • WorkOS – everything you need to make your app enterprise ready. Two things I found especially interesting, talking with Hillel: 1. Amazon used TLA+ to find a bug almost impossible to locate without formal methods. In the paper How AWS uses formal methods, the AWS team shared that they’d found a complicated bug for which the shortest error trace to exhibit was 35 steps (!!). The bug passed unnoticed through extensive design review, code reviews, and testing. AWS concluded they wouldn’t have uncovered it if they’d stuck to conventional testing approaches. 2. Why not use formal verification for everything, then? It’s because specs in the real world are a nightmare to write. Even a simple problem like “find the file in a directory that has the most lines” gets complicated when modeled with formal methods. We would have to answer questions like: ‘do we look at ASCII or UTF-8 new line characters, what about unreadable files, and Symlinks?’ Without formal methods, we can write a simple verification that is right in 99%+ of cases. Formal methods require a lot of extra effort for the less than 1% of exotic use cases!
Gergely Orosz34,525 views • 1 month ago

Chris Lattner (Chris Lattner) is one of the most influential engineers of the past two decades. He created LLVM, Swift, contributed to TensorFlow, and created the Mojo programming language. What was the story about creating Swift - and why did he face resistance inside Apple when wanting to replace Objective C? What did he learn at Tesla, Google and CPU maker SiFive, that led him to working on Mojo at Modular? We cover these and many more in today's episode. Watch or listen: • YouTube: • Spotify: • Apple: Brought to you by: • Statsig — The unified platform for flags, analytics, experiments, and more. • Linear – The system for modern product development. My favorite quote from Chris in this episode: “I believe in the power of programmers. I believe in the human potential of people that want to create things. And that’s fundamentally why I love software is that you can create anything that you can imagine.”
Gergely Orosz206,271 views • 10 months ago

How did a tiny team of 30 engineers build WhatsApp, more than a decade ago? From Jean Lee, engineer #19 at the company. Timestamps: 00:00 Intro 01:39 Early years in tech 06:18 Becoming engineer #19 at WhatsApp 13:53 WhatsApp’s tech stack 18:09 WhatsApp’s unique ways of working 25:27 Countdown displays and outages 27:07 Why WhatsApp won 28:53 The Facebook acquisition 33:13 Life after acquisition 39:27 Working at Facebook in London 44:07 Transitioning to management 47:27 Performance reviews as a manager 53:29 After Facebook 58:53 AI’s impact on engineering 1:02:34 Jean’s advice to new grads and startups 1:06:45 Empowering employees 1:08:17 Book recommendations Watch or listen: • YouTube: • Spotify: • Apple: Brought to you by: • Statsig – The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for automated code review. • WorkOS – Everything you need to make your app enterprise ready Three interesting observations from this episode: 1. WhatsApp had no code reviews after in-place. WhatsApp cofounder, Brian Acton, reviewed the very first pull request of each new hire, and after that, there were no more code reviews. Jean recounts how Brian reviewed her debut PR in extreme detail. This first (and only!) review set the bar high, and she wrote code to that standard from then on. 2. WhatsApp had close to zero formal processes. WhatsApp had no Scrum, no Agile, no TDD (test driven development), and no formal code reviews beyond the first commit. In contrast, Skype had 1,000 engineers and mandatory Scrum training, but WhatsApp still outcompeted it and won. Jean’s response to hearing of all the formal processes Skype used in order to execute faster: “I’m surprised to hear they thought they were shipping faster because of it.” Perhaps process is often a substitute for trust, not quality?” 3. Saying “no” to features was a competitive advantage. WhatsApp’s CEO, Jan Koum, rejected 99% of feature requests from the team. While competitors shipped dozens of shiny, new features, WhatsApp ruthlessly prioritized reliability and simplicity. Jan repeatedly told the team what the mission was. “I want a grandma living in the countryside to be able to use our app”, he said.
Gergely Orosz111,994 views • 5 months ago

How is the world's largest storage system, AWS S3 built? I found the best person to talk about this: Mai-Lan Tomsen Bukovec, who has been heading up S3 for 13 years (!) S3's scale is something else (tens of millions of hard drives, eleven 9s of durability (!!) and heavy usage of formal methods, microservices dedicated to durability, amongst others.) Watch or listen: • YouTube: • Spotify: • Apple: Brought to you by: • Statsig — The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for automated code review. Check out their latest State of Code Developer Survey report: • WorkOS – Everything you need to make your app enterprise ready.
Gergely Orosz133,719 views • 7 months ago

Kelsey Hightower has one of the most inspiring stories in tech: he went from a technician installing DSL modems, through self-directed study and very hard work, to one of the very few Distinguished Engineer at Google whom Satya Nadella personally persuaded to join Microsoft. Timestamps: 00:00 Intro 03:34 Kelsey’s first job at McDonald’s 05:04 His non-traditional path into tech 11:45 Landing his first tech job with an A+ certification 15:33 His entrepreneurial years 19:45 Joining Google as a data center technician 27:48 Learning automation at a Rackspace spinoff 33:26 Moving into financial services 50:00 Building a reputation through open source 53:55 From configuration management to containers 1:08:20 The rise of Kubernetes 1:25:05 Why he almost joined NASA instead of Google 1:29:20 Defining DevRel at Google 1:38:20 Demonstrating impact at Google 1:41:20 Microsoft's offer 1:55:20 Learning how to slow down 2:06:39 Advising and investing 2:15:03 A people-first view of GenAI 2:24:27 Using AI with guardrails 2:28:26 Matching AI to the task 2:36:06 Staying relevant in the AI era Brought to you by outstanding teams building products I love: • Antithesis: verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. • Sentry: application monitoring software considered “not bad” by millions of developers • Buildkite: CI software built to absorb whatever your coding agents throw at the build queue. OpenAI, Anthropic, Uber and others are customers: Three interesting learnings from Kelsey: 1. Side hustles and doing your own thing teach you business like no IC job can. Before becoming a software engineer at Google, Kelsey was a manager for his comedian friend, operated a computer store, and did IT contracting. These gigs taught him logistics, planning, and about money. All this helped him be far more effective at talking with executives and acting as an executive sponsor inside Google. 2. Can you explain what your startup does without mentioning AI? When Kelsey researches startups seeking his advice, he challenges founders to not say “AI” once. This means that they must explain the actual value their company creates. One unexpected benefit of this is that it often reveals there are easier, cheaper ways to achieve a goal than with AI. 3. It’s very rare to get an extra zero put on your compensation figure – but it happened. Kelsey was a successful, well-paid Google engineer when Microsoft made him an offer that 10x’d his salary (!!). When Kelsey told Google he was planning to take the offer, it matched the offer, proving that his market value had massively increased. It shows that being well paid doesn’t necessarily mean you’re being paid at the correct market rate.
Gergely Orosz61,547 views • 3 months ago

How have the fundamentals of building large, distributed software systems changed the last decade? A conversation with Martin Kleppmann (author of Designing Data-Intensive Applications) - given that the second, updated edition of the book was just released. Timestamps: 00:00 Early career 05:46 Building Rapportive 10:47 Working at LinkedIn 14:09 Writing Designing Data-Intensive Applications 23:00 Reliability, scalability, and repeatability 26:24 DDIA: the second edition 30:50 Tradeoffs of using cloud services 39:02 How the cloud changed scaling 42:53 The trouble with distributed systems 49:02 Ethics for software engineers 52:45 Formal verification 1:00:12 Academia vs. industry 1:03:50 Local-first software 1:09:50 Computer science education 1:18:32 Martin’s current research and advice Brought to you by: • Statsig – The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for code verification and automated code review. Check out Sonar's new architecture management capabilities that ensure both humans and AI agents respect your system’s blueprint. • WorkOS – Ship enterprise features – SSO, directory sync, RBAC, audit logs – in days, not months. Three things worth considering, as discussed with Martin, in this episode: 1. Multi-region and multi-cloud are risk/cost trade-offs, not best practices. Martin does not believe that there is a “best practice” in deciding whether to go multi-region or multi-cloud. This decision is a tradeoff between risk and costs. It’s a business decision to be made. Designing Data-Intensive Applications gives engineers the vocabulary to articulate the tradeoffs, not to dictate answers. 2. Replication for fault tolerance is more relevant for most engineers these days than sharding. Though the book has a full chapter on sharding, Martin said that the cloud has reduced the need for manual sharding for the majority of teams. This is also because machines are increasingly bigger, and more workloads fit on a single machine. Sharding across machines is increasingly a specialist concern; replication for fault tolerance, however, is still relevant at every scale. 3. Knowing system internals as a superpower for application developers. Martin maintains that Designing Data-Intensive Applications is not a book for people who build databases or even infrastructure, but it’s helpful for application developers to develop an intuition for making good design decisions and debugging performance issues we will eventually encounter.
Gergely Orosz79,406 views • 4 months ago

Why is Rust different than many/most programming languages? Alice Ryhl works on Google's Android Rust team, is a Rust language team advisor, and is a core maintainer of Tokio (the most widely-used async runtime in Rust) Timestamps: 00:00 Intro 04:09 Tokio: an overview 05:11 What Alice likes about Rust 12:48 Rust for TypeScript engineers 13:51 Moving from C++ to Rust 14:34 Memory safety 18:12 Garbage collection tradeoffs 21:46 Ownership, references, and borrowing 26:59 Unsafe in Rust 31:21 Crates and Cargo 35:55 Language design and RFCs 43:02 Building new features 46:30 Editions vs. versions 49:47 Getting paid to work on Rust 51:27 Contributing to Rust 53:03 Rust in the Linux kernel 55:45 AI use cases for Rust 1:01:35 Learning Rust 1:03:54 Book recommendation Brought to you by: • Antithesis – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. • Sentry – application monitoring software considered “not bad” by millions of developers Three things worth knowing about Rust: 1. Rust was designed to turn implicit failures into compile errors. Where other languages allow you to forget something, Rust makes an omission into a compilation error for things like null checks, uninitialized variables, or error propagation with the ‘?’ character. If you mess something up, it’s almost certain your program will not compile. If it does, at the very least you should see a lint warning. 2. Refactoring in Rust is safe and easy, thanks to the compiler. Alice: “I change a return type or struct field, then just fix the compiler errors until the compiler stops shouting. And then once I’ve done that, I’ve updated every place I need to update.” Rust’s focus on correctness makes refactoring it more straightforward than dynamically-typed languages and Java-style typed ones are to refactor. 3. “Editions” allow Rust to make breaking changes without ‘breaking’ anyone’s code. Rust editions (2015, 2018, 2021, 2024) can be mixed freely across crates. A library on the 2021 edition works seamlessly with a binary on the 2024 edition. This is how Rust evolves syntax (like adding async/await as keywords) without forcing an ecosystem-wide migration. Thanks a lot, Alice for this great discussion! And for your work on Rust.
Gergely Orosz51,961 views • 3 months ago

Why did Uber build thousands of microservices? No better person to answer than Uber's first CTO, Thuan Pham. Timestamps: 00:00 Intro 05:32 Getting into tech 16:09 The dot-com bust 20:42 VMware 26:29 Getting hired by Travis at Uber 33:22 Early days at Uber and scaling challenges 40:57 Uber’s China launch 47:12 The platform and program split 50:26 From monolith to microservices 53:38 Internal tools at Uber 57:05 Helix: Uber’s mobile app rewrite 59:55 Thuan’s email about naming 1:02:03 Org structure changes under 1:06:34 Thuan’s work philosophy 1:12:23 The “three tours of duty” at Uber 1:15:37 Why Thuan left Uber 1:17:34 Coupang and Nubank 1:21:59 Faire 1:25:31 How Faire uses AI 1:28:24 AI’s impact on software engineering 1:31:09 The role of the CTO 1:35:13 Career advice Brought to you by: • Statsig – The unified platform for flags, analytics, experiments, and more. • WorkOS – Everything you need to make your app enterprise ready. • Sonar – The makers of SonarQube, the industry standard for automated code review. Check out SonarQube Advanced Security: Three interesting parts from this conversation: 1. The program/platform split came before microservices. The concept of cross-functional “program” teams and dedicated “platform” teams became necessary because an org split across backend, frontend and mobile engineers slowed down in execution speed when Uber grew to around 100 engineers. Every feature required negotiating bandwidth across the mobile, backend, and dispatch teams. Thuan, Travis Kalanick, and Jeff Holden literally used color-coded sticky notes with people’s names to reorganize into self-sufficient teams. We cover more about this split in this The Pragmatic Engineer deepdive, The Platform and Program split at Uber: 2. Expect multiple rewrites during hypergrowth. The right architecture depends on how fast a product and company are growing. At Uber, repeated rewrites were common because each one “bought” another window of survival for the company. Thuan’s recommendation is to understand that a rewrite simply means a company is outrunning its existing architecture: this is not necessarily a bad thing! 3. Uber is the only major company that had a “Senior 1” and “Senior 2” level – and Thuan is unapologetic. Thuan introduced the Senior 1 (L5A) and Senior 2 (L5B) levels because the jump from senior (L5) to Staff (L6) became very big, and larger than between previous levels. One problem this split level created was that Uber’s L5B was akin to Google’s and Facebook’s L6/E6. Thuan resisted the title inflation of just renaming L5B to ‘Staff’.
Gergely Orosz69,369 views • 5 months ago

What if we're actually in the middle of the third golden age of software engineering? This is what Grady Booch sees happening. If you are anxious about the state of the industry, you want to watch/listen to Grady's longer-term perspective and stories. Watch the full episode here: 00:00 Intro 01:58 The first golden age of software engineering 18:59 The software crisis 33:01 The second golden age of software engineering 42:21 Y2K and the Dotcom crash 45:47 Early AI 47:34 The third golden age of software engineering 51:48 Why software engineers will very much be needed 58:46 Grady responds to Dario Amodei 1:06:54 New skills engineers will need to succeed 1:10:04 Resources for studying complex systems 1:14:33 How to thrive during periods of change Brought to you by: • Statsig — The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for automated code review. Join me online at the Sonar Summit on March 3rd, where I talk about practical tactics for the AI era. • WorkOS – Everything you need to make your app enterprise-ready.
Gergely Orosz87,027 views • 7 months ago