Blog
System design, taken apart properly.
Long-form breakdowns of real systems — the architecture, the failure modes, and how to actually present them in an interview.
78 articles across 8 topics.
System Design
25 articles
Stop Adding AI to Everything: Choose the Simplest Reliable Solution
Log alerts, rate limits, slow queries, caches, service calls, and deployments: solve the deterministic problem first. Add AI where uncertainty actually exists.

Little's Law Explained: The One Formula That Connects Latency, Throughput, and Concurrency
Little's Law says concurrency equals arrival rate times time-in-system. One line of arithmetic sizes your thread pools, connection pools, and load tests — and explains why latency and throughput are the same knob.

Behind the Scenes: 11 Technologies You Use Every Day
Follow the real paths behind eleven familiar technologies—from a B+ tree page lookup and JWT verification to DNS resolution, a TLS handshake and health-aware load balancing.

Tech You Know But Don't Really Understand: What Happens Underneath
Knowing the name isn't understanding the technology. Ten everyday systems — Kafka, Redis, Kubernetes, indexes, load balancers, CDNs, Docker, API gateways and WebSockets — explained by what happens underneath.

Distributed Locks: What They Are and How Not to Get Burned
Two workers, one cron, one payout — sent twice. A mutex cannot help you here: there is no shared memory to guard. This is what a distributed lock really buys you, and what it quietly does not.

How to Design ChatGPT at Scale
The model is not the interesting part. The interesting part is that every request occupies expensive hardware for many seconds, streams its output, and costs real money per token.

How to Start ANY System Design Interview
Most system design interviews are decided in the first ten minutes. Not by what you design, but by whether you established what you were designing for.

Design a Rate Limiter for 10M Users
Choosing the algorithm is the easy part. The interview is really about distributed shared state: how twenty instances agree on one counter without adding a round trip to every request.

Users Are Getting Duplicate Payments. Find the Bug.
Duplicate charges are almost never a UI problem. They are a retry meeting a write that was never idempotent — and the fix is a unique constraint written before the money moves.

How Blinkit Finds the Nearest Delivery Partner
The nearest-rider query is the easy half. The hard half is that every rider reports a new location every four seconds, and the answer must be correct three seconds later.

Your Kafka Consumer Is Getting Slower. Now What?
Consumer lag has three completely different causes that look identical on a dashboard. Aggregate lag hides all of them — per-partition lag tells you which one you have.

Your Redis Cache Suddenly Became Useless
Redis is almost never the problem. The caching strategy wrapped around it is. Here is how to find out which of the seven failure modes you are in, using numbers Redis already gives you.

How Instagram Handles 1 Billion Likes
A like looks trivial and is one of the best scale questions there is: extreme write volume, extreme read fan-out, hot keys, and a count that must be fast but only approximately correct.

What Happens When 10M Users Hit Your API?
Traffic does not break systems uniformly. It finds the one resource with the lowest ceiling — usually database connections or worker threads — and everything downstream cascades from there.

Your Database Is at 100% CPU. What Do You Do?
A database at 100% CPU is rarely out of capacity. It is usually one bad plan, one missing index, or one broken cache. Here is the order to check things in — and how to say it in an interview.

Top 10 System Design Interview Questions (With Answers)
Ten questions cover most of what senior system design interviews actually test. Here is a crisp answer to each, the trap hidden inside it, and a deep dive link when you want the full reasoning.

Firewall Explained: How It Actually Secures an Organisation
A firewall is not just a box that blocks ports. It is a policy engine that decides which traffic is trustworthy based on identity, state, behaviour and context. Here is how that decision is made.

How MCP Works: The Model Context Protocol, Explained by Design
MCP is not a product, it is a wire protocol. Once you see it as JSON-RPC plus a capability handshake plus two transports, the whole thing becomes obvious — and so do its failure modes.

Graph Engineering in AI: Agentic Graphs vs Knowledge Graphs
One phrase, two disciplines. Agentic graph engineering shapes how an LLM application executes; knowledge graph engineering shapes what it knows. Confusing them is why so many AI architecture discussions go sideways.

Build Your Own Knowledge Graph: A Practical End-to-End Guide
Most knowledge graph tutorials stop at 'here is a node and here is an edge'. This one walks the whole pipeline: ontology first, extraction second, entity resolution third, and only then the database.

Build Agentic Workflows with Graph Engineering (LangGraph in Practice)
A prompt chain is a straight line. Real agents need loops, retries, human approval and resumability — and that is exactly what modelling the workflow as a graph gives you.

GraphRAG vs Vector RAG: Which Retrieval Architecture Actually Fits
Vector search answers 'what does the corpus say about X'. It cannot answer 'which suppliers are two hops from a sanctioned entity'. Knowing which question you have is the whole architecture decision.

Entity Resolution: The Hardest Part of Any Knowledge Graph
Every failed knowledge graph project I have seen failed here. Extraction is easy now; deciding that 'Acme Corp' and 'ACME Corporation Ltd' are one node is still genuinely hard engineering.

Multi-Agent Orchestration Patterns: Supervisor, Pipeline, Swarm and Hierarchy
Multi-agent is a topology decision, not a vibe. Pipeline, supervisor, hierarchy and swarm each fail differently — pick by coupling and failure containment, not by how impressive the diagram looks.

How UPI Payments Actually Work — NPCI, System Design & the Interview Playbook
13+ billion transactions a month, ~7,000 TPS at peak, sub-second settlement, and near-zero tolerance for money loss. Here's how UPI is built — and how to design it on a whiteboard.
Architecture
14 articles
Kafka Explained: Partitions, Offsets, Delivery Guarantees and Rebalances
Kafka is a distributed, append-only log you can replay. Once you accept that ordering only exists inside a partition, almost every design question answers itself.

Hexagonal Architecture (Ports and Adapters): Dependencies Point Inward
Your business rules should not import the web framework, the ORM, or the payment SDK. Invert the arrows and the domain becomes the most stable, fastest-testing part of the system.

Pipe-and-Filter Architecture: Composable Transformation at Any Scale
Each stage knows its input contract and its output contract, and nothing else. That single constraint is what makes pipelines parallelisable, testable and endlessly recomposable.

Space-Based Architecture: Removing the Database From the Hot Path
You can add replicas until the servers are free. If every request still ends at one database, the ceiling never moves. Space-based architecture moves the ceiling.

Peer-to-Peer Architecture: Systems With No Centre
Remove the server and capacity grows with users, failure has no single point — and you inherit discovery, NAT traversal and the problem that peers can lie.

Serverless Architecture: Scale to Zero, and the Bill for It
Serverless is not 'no servers'. It is 'someone else's servers, billed per millisecond, with constraints you must design around'.

Microkernel (Plugin) Architecture: A Small Core and Infinite Variation
When the requirement is 'the same product, endlessly customised', the answer is almost never more if-statements. It is an extension contract.

Service-Oriented Architecture (SOA): The Enterprise Ancestor of Microservices
Microservices did not replace SOA; they dropped the bus and shrank the services. Understanding what SOA got right — and what the ESB got wrong — explains most of modern service design.

Microservices Architecture: What You Actually Buy and What You Pay
Microservices are an organisational technology with technical consequences. If your teams are not blocked on each other, you are paying the bill without buying anything.

Client-Server Architecture: Drawing the Trust Boundary
Thin client, thick client, offline-first, optimistic UI — these are all the same question: where does authority live, and what happens when the network disagrees?

Layered (N-Tier) Architecture: The Industry Default, Done Properly
Almost every codebase you will ever join is layered. Almost none of them chose it deliberately, which is why so many of them hurt.

The Modular Monolith: Microservice Boundaries Without the Bill
The hard part of microservices is the boundaries, not the network. A modular monolith makes you do the hard part first — and lets you skip the network until it earns its place.

Monolithic Architecture: Still the Right Default
Most teams that regret their architecture did not regret the monolith. They regretted the twelve services they built before they understood the domain.

13 Software Architectures Every Engineer Should Know
Architecture is not a menu of trends. Each style is an answer to a specific pressure — deployment speed, blast radius, team size, latency, cost. Here are thirteen styles and the pressure each one solves.
Design Patterns
11 articles
Event-Driven Architecture: Loose Coupling and the Bill That Comes With It
Adding an analytics consumer should not require a pull request on the order service. That is the promise. The price is that 'done' no longer has a single moment in time.

Strangler Fig: Replacing a Monolith Without a Big-Bang Rewrite
Two years of parallel development, a frozen legacy system and a launch weekend nobody survives. Or: move /products this sprint, keep everything else running, and repeat.

API Gateway: One Front Door for Many Services
Ten services, three clients, thirty integration points and auth implemented ten slightly different ways. The gateway collapses that into one edge you can actually reason about.

Cache-Aside: The Default Caching Strategy and Its Sharp Edges
Your database should not answer the same query ten thousand times an hour. Caching is easy; invalidation, stampedes and staleness are the actual engineering.

CQRS: When Reads and Writes Stop Wanting the Same Database
Your normalised order schema is perfect for writes and terrible for the six-join query powering search. CQRS stops pretending one model can be optimal for both.

Idempotency: Why the Same Request Twice Must Charge Once
The user tapped PAY NOW once. The network timed out, the SDK retried, and ₹5,000 left their account twice. Idempotency is the cheapest insurance in distributed systems.

Transactional Outbox: When the Database Commits but Kafka Never Hears
The order exists. The event does not. Inventory never reserved, the warehouse never printed a label, and no error was logged anywhere. The dual-write problem, and its standard fix.

Bulkhead Pattern: Containing the Blast Radius
Recommendations went slow and payments stopped. Not because they were related — because they shared a thread pool. Bulkheads make that impossible by construction.

Circuit Breaker: Stop One Slow Service From Taking Down Everything
Checkout did not break. It waited. Threads piled up behind a slow recommendation call until the pool was gone. A circuit breaker is how you refuse to wait.

Saga Pattern: Transactions That Span Multiple Services
Payment succeeded, inventory failed, the customer is charged for stock you do not have. A saga is the disciplined answer to partial failure across service boundaries.

10 Design Patterns You Need in the AI Era
Code generation got cheap. Reasoning about production failure modes did not. These ten patterns are the vocabulary senior engineers use to describe what breaks and why it stays contained.
APIs & HTTP
10 articles
Realtime APIs: WebSockets vs SSE vs Long Polling vs Webhooks
Most teams reach for WebSockets when Server-Sent Events would do the job over ordinary HTTP with free reconnection. Here is how to tell the difference before you commit.

gRPC Explained: Protobuf, HTTP/2, Streaming and Deadlines
Contract-first RPC over HTTP/2 with binary payloads. Smaller, faster and strictly typed — at the cost of browser support and human readability.

GraphQL Explained: Schema, Resolvers, N+1 and When Not To Use It
One endpoint, a typed schema, and clients that ask for exactly what they need. The power is real; so are the N+1 queries, the lost HTTP caching and the DoS surface.

REST API Design: Constraints, Resources, Versioning and Pagination
REST is a set of architectural constraints, not a JSON-over-HTTP aesthetic. The constraints are what buy you caching, scalability and clients that do not break.

Types of APIs Explained: REST, GraphQL, gRPC, WebSockets, Webhooks and More
There is no best API style. There is a request/response axis, a streaming axis, and a who-calls-whom axis — and most real systems use three of these at once.

HTTP Status Codes and Headers: The Practical Guide
Returning 200 with an error inside the body is the most common API design mistake there is. Status codes are the machine-readable half of your contract.

PUT vs PATCH vs POST: Choosing the Right Write Method
They all write. They differ in who owns the resource identity, whether the payload is complete, and what happens when the client retries after a timeout.

All HTTP Methods Explained: GET, POST, PUT, PATCH, DELETE and the Rest
Methods are contracts, not labels. Safety and idempotency decide what a proxy may cache, what a client may retry, and what a crawler is allowed to touch.

Everything You Need to Know About an HTTP Request
An HTTP request is a plain text envelope on top of a very carefully engineered stack. Here is the whole journey — resolution, handshake, framing, headers, body, response, reuse.

Every Part of a URL, Explained — Scheme, Host, Path, Query & Fragment
You read hundreds of URLs a day and probably can't name half their parts. Here's the full anatomy of a URL, with diagrams — and the rule for choosing path parameters over query parameters in API design.
DevOps & Cloud
8 articles
5 DevOps + AI + Cloud Projects That Actually Stand Out
Not another CRUD app. Five projects—incident copilot, AI CI/CD, self-healing Kubernetes, cost optimizer and security engine—with build guides and reference repos.

5 DevOps Concepts Developers Usually Get Wrong
Five familiar DevOps terms hide important operational differences. Learn the mechanism, the common mistake, and what changes in production.

How Code Flows: From Local Development to Production
A visual, end-to-end trace of how one code change becomes a versioned production deployment—and how the platform validates, routes, observes, scales and rolls it back.

Nginx Explained: Reverse Proxy, Load Balancing, TLS and Tuning
Nginx sits in front of almost everything you deploy. Understanding its worker model, its upstream timeouts and its caching rules is the difference between a 502 you can explain and one you cannot.

The Request That Never Reached Your Server
A 200 OK does not mean your backend generated the response. Between your browser and your server there are five-plus layers, and the request can stop at almost every one of them.

What Is AWS? A Developer's Guide to the Basics
AWS has hundreds of services, but you only need about fifteen to build almost anything. Here is what AWS actually is, how regions and availability zones work, the core building blocks, and how billing and security really function.

The AWS Trillion-Dollar Billing Bug: What Happened and What Developers Should Learn
A configuration change put a unit-pricing error into the AWS estimated-billing pipeline. Alarms fired but paged nobody, the rollback failed, and cost alerts were switched off platform-wide. The absurd numbers are the least interesting part.

SQL Injection Explained: How It Works, When It Happens, and How to Stop It
SQL injection is still one of the most common ways applications leak or destroy data. Here is how the attack works, what vulnerable code looks like, and how to defend against it in production.
Databases & Search
3 articles
Inverted Indexing Explained: From Documents to Millisecond Search
Search engines do not scan documents at query time. They reverse the problem: terms point to documents. Follow one sentence through analysis, postings, intersections, scoring, segments and distributed retrieval.

Elasticsearch Explained: Inverted Indexes, Sharding, Relevance and Sizing
Search feels like magic until you realise it is one data structure — the inverted index — plus a distributed system wrapped around it, with a refresh interval that quietly defines your consistency model.

Redis Explained: Data Structures, Persistence, Clustering and Real Use Cases
Redis is not just a cache. It is an in-memory data-structure server with a single-threaded core, and almost every production problem with it comes from misunderstanding that one sentence.
Frontend
3 articles
Debouncing vs Throttling: Which One to Use and Why
Both limit how often a function runs, but they answer different questions. Debounce asks 'has it stopped?', throttle asks 'has enough time passed?'. Here is how to pick the right one every time.

Throttling Explained: Rate-Limiting Work Without Dropping It
Throttling guarantees a function runs at most once per interval, no matter how many events arrive. Here is how to implement it correctly on the client, and how the same idea becomes rate limiting on the server.

Debouncing Explained: How It Works and When to Use It
Debouncing collapses a burst of events into a single call once the burst stops. Here is the mental model, the leading/trailing variants, production-grade implementations, and the traps that break it.
Careers & Craft
4 articles
Forward Deployed Engineer (FDE) Roadmap: The Role, The Loop, The Prep
Half product engineer, half consultant, entirely accountable for the customer's outcome. A practical roadmap for the fastest-growing engineering title in AI.

100+ Engineering Blogs Every Developer Should Read (2026 Edition)
Independent engineering blogs are where mental models come from; company blogs are where case studies come from. Here is the full map — 12 categories, 100+ links, plus the 20 that matter most if you only have time for a shortlist.

12 Free Tools That Make Developers' Lives Easier
Every tool here solves a specific daily problem: sketching architecture, decoding a JWT, testing an API, reading docs offline, or finding out what your image files are quietly carrying. All free, most open source.

How Your Phone Knows You're Walking, Running or Driving
Your phone silently classifies you as still, walking, running, cycling or driving — all day, on a few milliwatts. Here is the actual pipeline: raw IMU signals, gravity separation, windowed features, a classifier, and a smoothing layer that hides its mistakes.