05 — 51 patterns · 12 stacks · 52 books
Library
The knowledge that builds a capability: patterns in our own words with one canonical reference each, reference stacks, and the books worth an architect's week.
// FIG. 05 · Design patterns
Resilience · 5
- DP-010
Circuit Breaker
Wrap a call to a remote dependency in a switch that opens after repeated failures, so callers fail fast and the struggling dependency gets room to recover instead of a stampede.
When · When a slow or failing downstream service would otherwise tie up threads and cascade; costs a fallback path and tuned thresholds.
Reference ↗related DP-020
- DP-020
Retry with Backoff
Repeat a failed call a bounded number of times with growing, jittered delays, so transient faults heal themselves without synchronized retries making them worse.
When · Only for idempotent operations and transient errors; pair with a circuit breaker and a deadline, or retries become a self-inflicted outage.
Reference ↗related DP-010
- DP-030
Bulkhead
Partition a system's resources (thread pools, connections, instances) by consumer or dependency, so one misbehaving tenant or slow downstream exhausts only its own compartment and the rest keep serving.
When · When one noisy workload or dependency can starve everything else; costs idle capacity in each compartment and a sizing decision per partition.
Reference ↗related DP-010DP-040
- DP-040
Rate Limiting
Cap how many requests a caller may make in a window and reject or queue the excess with a clear signal, so a shared service protects its own capacity and treats tenants fairly.
When · Any API exposed to many clients or to the public; costs a policy for who gets how much, and clients that know how to back off when told.
Reference ↗related DP-020DP-050DP-110
- DP-050
Throttling
Let a service shed or degrade lower-priority work when it approaches its limits, instead of accepting every request and failing all of them at once.
When · When demand can exceed capacity faster than it can scale; costs an explicit ranking of what to drop first and a degraded mode that still makes sense to users.
Reference ↗related DP-040DP-060
Scalability · 5
- DP-060
Queue-Based Load Leveling
Put a queue between producers and a slower consumer so bursts are absorbed and the consumer runs at a steady pace it can sustain, rather than being sized for the peak.
When · When arrival rate is spiky but work can tolerate a short delay; costs a queue to operate and an answer for what happens when the backlog grows faster than it drains.
Reference ↗related DP-050DP-100
- DP-070
Cache-Aside
The application checks a cache first, loads from the system of record on a miss, and writes the result back into the cache, keeping the cache a disposable copy rather than a second source of truth.
When · Read-heavy data that tolerates brief staleness; costs an invalidation strategy, and gets dangerous when the cache quietly becomes the only place a value lives.
Reference ↗related DP-090
- DP-080
Sharding
Split one logical dataset across several stores by a key, so write volume, storage and hot spots are spread instead of concentrated in a single database.
When · When a single database can no longer hold the data or the write rate; costs cross-shard queries, rebalancing when a key gets hot, and a key choice that is hard to change later.
Reference ↗related DP-090DP-270
- DP-090
Read Replicas
Copy a primary database to one or more replicas that serve read traffic, so reporting and browse workloads stop competing with writes on the primary.
When · When reads dwarf writes and slightly stale results are acceptable; costs replication lag that the application must account for, especially right after a write.
Reference ↗related DP-070DP-080DP-270
- DP-100
Competing Consumers
Several identical workers pull from the same queue, so throughput scales by adding workers and a failed worker's message is simply picked up by another.
When · Independent units of work that can run in any order; costs idempotent processing and, where ordering matters, partitioning so related messages land on one consumer.
Reference ↗related DP-060DP-180
Integration · 9
- DP-110
API Gateway
A single entry point in front of many backend services that handles routing, authentication, rate limits and protocol translation, so each service does not reimplement the edge.
When · As soon as more than a handful of services face external clients; costs a component that can become a bottleneck or a dumping ground for business logic if left unguarded.
Reference ↗related DP-040DP-120DP-360
- DP-120
Backend for Frontend
Give each client type (web, mobile, partner) its own thin backend that shapes and aggregates data for that experience, instead of one general-purpose API that serves everyone poorly.
When · When clients differ enough that one API shape forces chatty calls or over-fetching; costs one more deployable per client and duplicated aggregation logic to keep honest.
Reference ↗related DP-110DP-150
- DP-130
Publish-Subscribe
Producers publish events to a topic without knowing who listens; any number of subscribers receive each event independently, so adding a consumer never requires changing the producer.
When · When several systems need to react to the same fact; costs event contracts that must be versioned, and the loss of a direct answer to the question "did anyone handle this".
Reference ↗related DP-140DP-230DP-310
- DP-140
Request-Reply
Carry a request over messaging and route the answer back on a reply channel keyed by a correlation id, getting a synchronous-style interaction with the buffering and decoupling of a queue.
When · When a caller needs an answer but the provider should not be called directly; costs correlation bookkeeping, timeouts and a plan for replies that never arrive.
Reference ↗related DP-130DP-150
- DP-150
Scatter-Gather
Broadcast one request to several providers in parallel, then combine their responses into one result, typically with a deadline so a slow participant does not hold the answer hostage.
When · Price checks, searches or availability lookups across many sources; costs aggregation rules for partial results and clear handling of the provider that never answers.
Reference ↗related DP-140DP-120
- DP-160
Content-Based Router
Inspect each message and send it to the right destination based on its content, so senders do not need to know which of several systems owns a given case.
When · When the same message type is handled by different systems depending on region, product line or customer segment; costs routing rules that need an owner and a test suite.
Reference ↗related DP-130DP-170
- DP-170
Claim Check
Store a large payload in shared storage and pass only a reference through the messaging system, so brokers stay fast and message size limits stop shaping the design.
When · Documents, images or bulk records moving through an event pipeline; costs a storage lifecycle for the payloads and a consumer that knows how to redeem the reference.
Reference ↗related DP-160DP-130
- DP-180
Idempotent Consumer
A consumer records the ids of messages it has already processed and ignores repeats, so at-least-once delivery does not create duplicate orders, payments or notifications.
When · Any consumer behind a broker that may redeliver; costs a deduplication store with a retention window and a message id that producers assign consistently.
Reference ↗related DP-100DP-290DP-020
- DP-190
Strangler Fig
Put a routing layer in front of a legacy system and move one capability at a time to a new implementation behind it, until the old system serves nothing and can be switched off.
When · Replacing a large system where a single cutover is too risky; costs running both systems for a long time and discipline to finish, since half-strangled systems are the usual outcome.
Reference ↗related DP-200DP-110
Domain Modeling · 6
- DP-200
Anti-Corruption Layer
A translation layer between your model and an external or legacy system's model, so foreign concepts and quirks are converted at the boundary instead of leaking into your core.
When · Integrating with a system whose model you do not control and do not want to inherit; costs a mapping layer that must evolve with both sides.
Reference ↗related DP-190DP-210
- DP-210
Bounded Context
Draw an explicit boundary inside which a domain model and its language are consistent, and accept that the same word ("customer", "order") can legitimately mean different things in different contexts.
When · When one enterprise-wide model keeps collapsing under conflicting definitions; costs explicit contracts and translation between contexts rather than one shared schema.
Reference ↗related DP-200DP-220DP-240
- DP-220
Aggregate
A cluster of domain objects treated as one unit for changes, with a single root that enforces its invariants, so consistency rules live in one place and transactions have a natural size.
When · When business rules span several objects that must change together; costs careful boundary sizing, since aggregates that are too large serialize everything and too small cannot enforce the rule.
Reference ↗related DP-210DP-230DP-260
- DP-230
Domain Events
Record things that happened in the business ("Order Placed", "Invoice Paid") as first-class, named, immutable facts, so other parts of the system react to them rather than polling state.
When · When side effects pile up inside commands or other teams need to know what happened; costs naming discipline and a decision about which events are internal and which are published contracts.
Reference ↗related DP-130DP-220DP-260
- DP-240
Modular Monolith
One deployable unit built from modules with enforced boundaries and explicit interfaces, keeping the operational simplicity of a monolith while preserving the option to extract a service later.
When · Small to mid-sized teams who want clean boundaries without distributed-systems overhead; costs tooling to enforce module boundaries, since a monolith with no fences is just a big ball.
Reference ↗related DP-210DP-250
- DP-250
Hexagonal Architecture
Keep the domain logic in the center and reach the outside world (databases, queues, UIs, other services) only through ports and adapters, so infrastructure can be swapped or faked without touching the core.
When · When business logic is worth testing in isolation and infrastructure choices may change; costs more interfaces and indirection than a simple CRUD application needs.
Reference ↗related DP-240DP-200
Data · 8
- DP-260
Event Sourcing
Persist every change as an appended event and rebuild current state by replaying them, so the full history is the source of truth and any past state can be reconstructed.
When · When audit, temporal queries or replay matter more than simplicity; costs event versioning, snapshotting for performance, and a learning curve the whole team pays.
Reference ↗related DP-270DP-230DP-220
- DP-270
CQRS
Separate the model that handles writes from the models that serve reads, so each can be shaped and scaled for its job instead of one schema compromising for both.
When · When read and write workloads diverge sharply or reads need many denormalized views; costs eventual consistency between the two sides and more moving parts to operate.
Reference ↗related DP-260DP-090DP-300
- DP-280
Saga
Run a business transaction across several services as a sequence of local transactions, each with a compensating action that undoes it if a later step fails.
When · When a workflow spans services that cannot share a database transaction; costs compensation logic for every step and acceptance that intermediate states are visible.
Reference ↗related DP-290DP-180DP-230
- DP-290
Transactional Outbox
Write the business change and the event to publish in the same database transaction, then have a relay publish from the outbox table, so state and messages never disagree.
When · Whenever a service both updates its database and emits events; costs a relay process and an outbox table to prune, which is still cheaper than a lost or phantom event.
Reference ↗related DP-300DP-180DP-280
- DP-300
Change Data Capture
Read a database's transaction log and stream each insert, update and delete as an event, so downstream systems follow the source without batch extracts or application changes.
When · Feeding analytics, caches or search from a system you cannot modify; costs log access on the source, schema-change handling, and events that describe rows rather than business meaning.
Reference ↗related DP-290DP-310DP-330
- DP-310
Schema Registry
Keep every event and message schema in a central, versioned registry with compatibility rules enforced at publish time, so a producer cannot silently break its consumers.
When · Any streaming platform with more than one team producing or consuming; costs a registry to run and compatibility rules that make some schema changes slower on purpose.
Reference ↗related DP-130DP-300
- DP-320
Data Mesh
Treat analytical data as products owned by the domains that know it, served on a shared self-service platform under federated governance, instead of one central team owning every pipeline.
When · Large organizations where the central data team is the bottleneck; costs real ownership in each domain and a platform good enough that teams do not each build their own.
Reference ↗related DP-330DP-210DP-480
- DP-330
Medallion Lakehouse Layering
Land raw data unchanged, refine it into cleaned and conformed tables, then publish curated, business-ready sets, so each layer has a known quality level and lineage back to the source.
When · Building an analytics platform on open table formats over object storage; costs storage for every layer and discipline about which layer consumers are allowed to read.
Reference ↗related DP-320DP-300
Security · 3
- DP-340
Zero Trust
Grant no access based on network location; every request is authenticated, authorized and encrypted against the identity of the user and device, with policy evaluated continuously rather than once at login.
When · Hybrid workforces, cloud estates and any system where the perimeter no longer exists; costs strong identity everywhere first, and a long migration from network-based controls.
Reference ↗related DP-360DP-430DP-350
- DP-350
Secrets Management
Keep credentials, keys and tokens in a dedicated vault that issues them at runtime, rotates them, and logs every access, so no secret lives in source code, config files or a chat thread.
When · From the first deployed service onward; costs a vault to operate and an application bootstrap that can authenticate to it without a secret of its own.
Reference ↗related DP-340DP-410
- DP-360
Federated Identity
Delegate authentication to a trusted identity provider and accept its signed tokens, so applications never handle passwords and users sign in once across many systems.
When · Any application with users who already have a corporate or social identity; costs a dependency on the provider's availability and careful validation of the tokens it issues.
Reference ↗related DP-340DP-110
Deployment · 3
- DP-370
Blue-Green Deployment
Run two identical production environments, deploy to the idle one, verify it, then switch traffic over in one move, keeping the old environment ready for an instant rollback.
When · Releases where downtime or a slow rollback is unacceptable; costs double the infrastructure during the switch and a plan for database changes that both versions must tolerate.
Reference ↗related DP-380DP-400
- DP-380
Canary Release
Send a small slice of real traffic to the new version, watch its error rate and latency against the old one, and widen the slice only while the signals stay healthy.
When · When production is the only environment that reveals the problem; costs traffic-splitting infrastructure, comparable metrics, and automated rollback when the canary fails.
Reference ↗related DP-370DP-390DP-470
- DP-390
Feature Flags
Separate deploying code from releasing a feature by wrapping it in a runtime switch, so unfinished work can ship dark and a bad feature can be turned off without a redeploy.
When · Trunk-based development and gradual rollouts; costs flag hygiene, since every flag left behind is a branch in production that nobody tests.
Reference ↗related DP-380DP-370
Cloud · 5
- DP-400
Immutable Infrastructure
Never modify a running server or image; every change builds a new artifact that replaces the old one, so what runs in production is exactly what was built and tested.
When · Containers, machine images and anything that can be rebuilt from source; costs a fast build pipeline and giving up the habit of fixing things by hand on the box.
Reference ↗related DP-410DP-370DP-440
- DP-410
Infrastructure as Code
Declare networks, compute, storage and permissions in version-controlled files and apply them with a tool, so environments are reproducible, reviewable and diffable like any other code.
When · Every cloud environment that will exist longer than a day; costs state management, a review process for infrastructure changes, and resisting the console when it is faster.
Reference ↗related DP-400DP-350
- DP-420
Sidecar
Deploy a helper process next to the application in the same unit, handling concerns like proxying, logging or secrets, so those capabilities arrive without being compiled into every service in every language.
When · Polyglot estates where a shared library is impractical; costs one more process per instance and a platform that understands the pairing.
Reference ↗related DP-430DP-450
- DP-430
Service Mesh
Move service-to-service networking (mutual TLS, retries, timeouts, traffic splitting, telemetry) out of application code and into a layer of proxies managed by a control plane.
When · Dozens of services where consistent security and traffic policy matter; costs operational complexity and latency that a small estate will never recoup.
Reference ↗related DP-420DP-340DP-380
- DP-440
Twelve-Factor App
A set of conventions for services that run well on cloud platforms: configuration from the environment, stateless processes, dependencies declared, logs as streams, disposable instances.
When · Any new service meant for a container or serverless platform; costs externalizing state and configuration, which legacy code written for a single server resists.
Reference ↗related DP-400DP-460DP-350
Observability · 3
- DP-450
Distributed Tracing
Propagate a trace id across every hop of a request and record timed spans in each service, so one slow user action can be followed through the whole system instead of guessed at from separate logs.
When · Any request that crosses more than two services; costs instrumentation in every service and a backend to store and query traces, with sampling once volume grows.
Reference ↗related DP-460DP-420
- DP-460
Structured Logging
Emit log events as machine-readable records with consistent fields (timestamp, level, service, trace id, tenant) rather than free text, so logs can be queried and correlated instead of grepped.
When · From the first service onward; costs agreeing on field names across teams and keeping sensitive values out of the fields people are most tempted to log.
Reference ↗related DP-450DP-440
- DP-470
SLO and Error Budget
Define a measurable target for user-facing reliability, and treat the gap between that target and perfection as a budget the team may spend on change; when it is exhausted, reliability work takes priority.
When · When "is it reliable enough" keeps being argued from anecdotes; costs honest measurement and leadership willing to honor the budget when a release date is at stake.
Reference ↗related DP-380DP-450
Organization · 4
- DP-480
Stream-Aligned Teams
Organize most teams around a continuous flow of work for one product or user journey, supported by platform, enabling and complicated-subsystem teams, so ownership and delivery sit in one place.
When · When handoffs between component teams dominate lead time; costs a real platform team and a reorganization that leadership must hold through the awkward period.
Reference ↗related DP-320DP-210DP-490
- DP-490
Inner Source
Run internal projects the way healthy open-source projects run: visible repositories, documented contribution paths, and maintainers who review pull requests from any team.
When · When shared components are blocked on one team's backlog; costs maintainer time for review and the cultural shift from "ask them" to "send a change".
Reference ↗related DP-480DP-500
- DP-500
Architecture Decision Records
Capture each significant decision as a short, dated document stating the context, the options, the choice and its consequences, kept next to the code so the reasoning survives the people who made it.
When · Every project with more than one architect or a lifespan over a year; costs a few paragraphs per decision and the habit of writing them before the decision calcifies.
Reference ↗related DP-510DP-490
- DP-510
Fitness Functions
Express architectural qualities (dependency direction, latency budgets, coupling limits) as automated checks that run in the pipeline, so the architecture is guarded continuously rather than reviewed once.
When · When review boards cannot keep up with the rate of change; costs choosing a few measurable qualities and accepting that the unmeasured ones still need a human.
Reference ↗related DP-500DP-240
// FIG. 06 · Reference stacks
- TS-010
Server-rendered React on serverless
Web applicationA React framework that renders on the server by default, deployed to a serverless edge platform, with a managed Postgres behind it. The default for a small team shipping a content-heavy product fast.
- Framework
- Next.js ↗
- Language
- TypeScript ↗
- Styling
- Tailwind CSS ↗
- Database
- PostgreSQL ↗
- Hosting
- Vercel ↗
Fits · Small teams, content and catalog products, anything where time-to-first-release matters more than control of the runtime.
patterns DP-020
- TS-020
Event-driven microservices on Kubernetes
Event streamingIndependently deployable services on a container orchestrator, talking through a durable event log rather than direct calls, with a mesh handling service-to-service security and traffic policy. The shape most large product organizations converge on once a monolith stops scaling with the team.
- Orchestration
- Kubernetes ↗
- Event log
- Apache Kafka ↗
- Service mesh
- Istio ↗
- Packaging
- Helm ↗
- Continuous delivery
- Argo CD ↗
- Service framework
- Spring Boot ↗
- Schema governance
- Confluent Schema Registry ↗
Fits · Organizations with many teams and real traffic; the platform cost only pays back when team autonomy and independent release are the bottleneck.
- TS-030
Open lakehouse analytics platform
Data platformRaw, refined and curated data as open-format tables on object storage, transformed by a distributed engine and a SQL modeling tool, scheduled by a workflow orchestrator. Analytics and machine learning read the same tables, which removes the warehouse-versus-lake split.
- Object storage
- Amazon S3 ↗
- Table format
- Apache Iceberg ↗
- Processing engine
- Apache Spark ↗
- SQL modeling
- dbt ↗
- Orchestration
- Apache Airflow ↗
- Query engine
- Trino ↗
- Change capture
- Debezium ↗
Fits · Data teams that want engine choice and open formats over a single vendor warehouse, and have the engineering depth to run the pieces.
- TS-040
Machine learning platform
Machine learningExperiment tracking, a model registry, a feature store and a serving layer on top of the container platform, so a model moves from a notebook to production through the same path every time and the lineage from training data to deployed version is recorded.
- Language
- Python ↗
- Training framework
- PyTorch ↗
- Experiment tracking and registry
- MLflow ↗
- Pipelines
- Kubeflow ↗
- Feature store
- Feast ↗
- Serving
- KServe ↗
- Orchestration
- Kubernetes ↗
Fits · Teams running more than a handful of models in production; a single model is better served by a managed endpoint than by this much platform.
- TS-050
Cross-platform mobile
MobileOne TypeScript codebase that ships native iOS and Android applications, with a managed build and over-the-air update service, so a web-fluent team can deliver mobile without two native teams.
- Framework
- React Native ↗
- Tooling and builds
- Expo ↗
- Language
- TypeScript ↗
- Navigation
- React Navigation ↗
- Release automation
- Fastlane ↗
- End-to-end testing
- Detox ↗
Fits · Product teams with React experience whose app is mostly forms, lists and content; graphics-heavy or hardware-bound apps still want native code.
- TS-060
Enterprise integration layer
API & integrationA gateway in front of published APIs, a routing and transformation engine for the system-to-system flows, a broker for asynchronous work, and a contract standard tying them together. The layer that lets a packaged-application estate change one system without renegotiating every interface.
- API gateway
- Kong Gateway ↗
- Routing and mediation
- Apache Camel ↗
- Message broker
- RabbitMQ ↗
- Event streaming
- Apache Kafka ↗
- API contracts
- OpenAPI Specification ↗
- Event contracts
- AsyncAPI ↗
Fits · Organizations with many packaged and legacy systems that must exchange data; the integration team becomes a product team or the layer becomes the new bottleneck.
- TS-070
Open observability stack
ObservabilityVendor-neutral instrumentation feeding separate stores for metrics, logs and traces, queried from one dashboard tool and alerted on by one rule engine. Owning the stack trades a license bill for operations work and data volume to manage.
- Instrumentation
- OpenTelemetry ↗
- Metrics
- Prometheus ↗
- Logs
- Grafana Loki ↗
- Traces
- Grafana Tempo ↗
- Dashboards
- Grafana ↗
- Alert routing
- Alertmanager ↗
Fits · Platform teams that can run stateful services and want to avoid per-host or per-GB pricing; small teams are usually better off with a hosted service.
- TS-080
Workforce identity stack
IdentityA central identity provider issuing standards-based tokens, a directory it authenticates against, automated account provisioning into applications, and a vault for machine credentials. Every application trusts the provider instead of keeping its own passwords.
- Identity provider
- Keycloak ↗
- Authentication protocol
- OpenID Connect ↗
- Authorization protocol
- OAuth 2.0 ↗
- Provisioning
- SCIM ↗
- Directory
- OpenLDAP ↗
- Machine secrets
- HashiCorp Vault ↗
Fits · Organizations standardizing sign-on across many applications; most will buy the provider as a service and keep the same shape.
- TS-090
Serverless API on AWS
API & integrationHTTP endpoints as individual functions behind a managed gateway, with a key-value store and a queue for asynchronous work, all declared as infrastructure code. No servers to patch and a bill that tracks requests, at the cost of cold starts and a vendor-shaped design.
- Compute
- AWS Lambda ↗
- Gateway
- Amazon API Gateway ↗
- Database
- Amazon DynamoDB ↗
- Queue
- Amazon SQS ↗
- Infrastructure as code
- Terraform ↗
- Language
- TypeScript ↗
Fits · Spiky or low-volume APIs, event handlers and integrations; steady high-throughput services usually cost less on containers.
- TS-100
Convention-driven monolith
Web applicationOne full-stack framework with strong conventions, a relational database, a cache, a background job runner and server-rendered pages enhanced with small amounts of JavaScript. One deployable, one team, most of the product built in a year.
- Framework
- Ruby on Rails ↗
- Database
- PostgreSQL ↗
- Cache and queues
- Redis ↗
- Background jobs
- Sidekiq ↗
- Front-end interactivity
- Hotwire ↗
Fits · Small teams building a business application end to end; the reliable choice when nobody can justify a distributed system yet.
- TS-110
Static site at the edge
Web applicationContent in version control, built to static HTML by a content-first framework and served from a global edge network, with a hosted pipeline rebuilding on every commit. Fast, cheap and nearly impossible to take down, as long as the site really is static.
- Framework
- Astro ↗
- Content
- Markdown ↗
- Hosting
- Cloudflare Pages ↗
- Build pipeline
- GitHub Actions ↗
- Styling
- Tailwind CSS ↗
Fits · Documentation, marketing and reference sites; the moment the site needs accounts or per-user content, look at TS-010 instead.
- TS-120
Platform infrastructure baseline
InfrastructureCloud accounts and networks declared as code, a managed container platform per environment, a pull-based delivery controller reconciling clusters from Git, and secrets and certificates issued automatically. The foundation an internal platform team offers before any product team deploys.
- Infrastructure as code
- Terraform ↗
- Container platform
- Kubernetes ↗
- GitOps delivery
- Argo CD ↗
- Secrets
- HashiCorp Vault ↗
- Certificates
- cert-manager ↗
- Policy
- Open Policy Agent ↗
Fits · Organizations running several product teams on shared cloud infrastructure; one team on one cluster can start with the managed provider's defaults.
// FIG. 07 · Books
Architecture fundamentals · 10
Fundamentals of Software Architecture
An Engineering Approach
Mark Richards, Neal Ford · 2020
The clearest modern statement that architecture is trade-off analysis, with a vocabulary of styles and characteristics you can use in a review the next day.
Software Architecture: The Hard Parts
Modern Trade-Off Analyses for Distributed Architectures
Neal Ford, Mark Richards, Pramod Sadalage, Zhamak Dehghani · 2021 · ISBN 9781492086895
Picks up where Fundamentals leaves off and works through the decisions that have no clean answer: where to cut a service, who owns the data, and how to document the trade-off you chose.
Clean Architecture
A Craftsman's Guide to Software Structure and Design
Robert C. Martin · 2017 · ISBN 9780134494166
The dependency rule in one book: business rules in the middle, frameworks and databases at the edge. Useful as a shared vocabulary even when a team does not adopt every layer.
patterns DP-250
A Philosophy of Software Design
John Ousterhout · 2018 · 2nd edition 2021
A short, opinionated case that complexity is the enemy and deep modules are the cure. The best book to hand an engineer who is about to design their first interface that others will depend on.
Patterns of Enterprise Application Architecture
Martin Fowler · 2002 · ISBN 9780321127426
The pattern names most enterprise codebases are built from, whether or not anyone read the book. Still the reference when a team argues about repositories, units of work, or where the domain logic lives.
Building Evolutionary Architectures
Automated Software Governance
Neal Ford, Rebecca Parsons, Patrick Kua, Pramod Sadalage · 2017 · 2nd edition 2022
Turns architecture governance from a review board into executable fitness functions. The practical answer to how a standard stays enforced after the architect leaves the room.
Software Architecture in Practice
Len Bass, Paul Clements, Rick Kazman · 1998 · 4th edition 2021
The textbook treatment of quality attributes and the tactics that achieve them. Dry, but it is where the discipline of stating a requirement as a measurable scenario comes from.
Software Engineering at Google
Lessons Learned from Programming Over Time
Titus Winters, Tom Manshreck, Hyrum Wright · 2020 · ISBN 9781492082798
Software engineering as programming integrated over time and people. Worth reading for the chapters on code review, dependency management, and deprecation, which apply well below Google's scale.
patterns DP-490
The Pragmatic Programmer
Your Journey to Mastery
David Thomas, Andrew Hunt · 1999 · 20th anniversary edition 2019 · ISBN 9780135957059
The habits that separate engineers who ship maintainable systems from those who do not, in short numbered tips. Architects quote it more than they admit.
Documenting Software Architectures
Views and Beyond
Paul Clements, Felix Bachmann, Len Bass, David Garlan, James Ivers, Reed Little, Paulo Merson, Robert Nord, Judith Stafford · 2002 · 2nd edition 2010
How to write an architecture down so a reader who was not in the meeting can use it: views, view packets, and what each stakeholder needs. Heavier than most teams want, but the view discipline carries over.
patterns DP-500
Distributed systems · 8
Release It!
Design and Deploy Production-Ready Software
Michael T. Nygard · 2007 · 2nd edition 2018
Where the stability patterns come from. Read it before the first production incident, not after.
Designing Data-Intensive Applications
The Big Ideas Behind Reliable, Scalable, and Maintainable Systems
Martin Kleppmann · 2017 · ISBN 9781449373320
The one book to read on how storage, replication, consensus, and stream processing actually behave under failure. It replaces a shelf of vendor documentation with the ideas underneath.
Building Microservices
Designing Fine-Grained Systems
Sam Newman · 2015 · 2nd edition 2021
The balanced account of what microservices cost as well as what they buy, with a strong emphasis on organizational fit. The second edition is the one to read.
Monolith to Microservices
Evolutionary Patterns to Transform Your Monolith
Sam Newman · 2019 · ISBN 9781492047841
Incremental decomposition without a big-bang rewrite: strangler fig, parallel run, and the hard problem of splitting a shared database. Read before any modernization business case.
Microservices Patterns
With Examples in Java
Chris Richardson · 2018 · ISBN 9781617294549
A pattern-language view of service decomposition, sagas, API composition, and transactional messaging. Most useful when a team needs a named option for a data-consistency problem.
Enterprise Integration Patterns
Designing, Building, and Deploying Messaging Solutions
Gregor Hohpe, Bobby Woolf · 2003 · ISBN 9780321200686
The vocabulary every integration platform borrows: channels, routers, translators, endpoints. Still the reference for reasoning about a message flow regardless of which broker carries it.
Designing Distributed Systems
Patterns and Paradigms for Scalable, Reliable Services
Brendan Burns · 2018
Container-level patterns (sidecar, ambassador, adapter) and the multi-node patterns built on them, written by one of the people who shaped Kubernetes. Short and concrete.
The Art of Scalability
Scalable Web Architecture, Processes, and Organizations for the Modern Enterprise
Martin L. Abbott, Michael T. Fisher · 2009 · 2nd edition 2015
Scalability as a problem of people and process as much as technology, with the scale cube as the model for splitting load. Dated in places, but the organizational chapters hold up.
Domain modeling · 4
Domain-Driven Design
Tackling Complexity in the Heart of Software
Eric Evans · 2003 · ISBN 9780321125217
The origin of bounded contexts, ubiquitous language, and aggregates. Read the strategic design chapters first; they matter more to an architect than the tactical patterns.
Implementing Domain-Driven Design
Vaughn Vernon · 2013 · ISBN 9780321834577
The companion that shows what Evans's ideas look like in code and in a team, including context mapping between systems that were never designed together.
Learning Domain-Driven Design
Aligning Software Architecture and Business Strategy
Vlad Khononov · 2021 · ISBN 9781098100131
The most approachable entry to DDD, with a clear link from subdomain type (core, supporting, generic) to the architecture and the level of investment each deserves.
Domain Modeling Made Functional
Tackle Software Complexity with Domain-Driven Design and F#
Scott Wlaschin · 2018
Shows how to make illegal states unrepresentable with types, so the domain model documents its own rules. The language is F#, but the modeling approach transfers to any typed language.
Data · 4
Data Mesh
Delivering Data-Driven Value at Scale
Zhamak Dehghani · 2022 · ISBN 9781492092391
The argument for domain-owned data products over a central lake, and what federated governance has to look like for that to work. Read critically; the operating-model cost is real.
patterns DP-320
Fundamentals of Data Engineering
Plan and Build Robust Data Systems
Joe Reis, Matt Housley · 2022 · ISBN 9781098108304
A tool-agnostic map of the data engineering lifecycle, from ingestion to serving, with the undercurrents (security, data management, orchestration) that every stage shares.
Streaming Systems
The What, Where, When, and How of Large-Scale Data Processing
Tyler Akidau, Slava Chernyak, Reuven Lax · 2018
Event time versus processing time, windows, triggers, and watermarks, explained once and properly. Required before designing anything that claims to be real-time.
Database Internals
A Deep Dive into How Distributed Data Systems Work
Alex Petrov · 2019 · ISBN 9781492040347
Storage engines in the first half, distributed consensus and replication in the second. The book to read when a database choice has to be defended on more than a benchmark.
Cloud & infrastructure · 4
Kubernetes Patterns
Reusable Elements for Designing Cloud Native Applications
Bilgin Ibryam, Roland Huss · 2019 · 2nd edition 2023
The patterns a platform team should standardize before the first workload lands: health probes, resource limits, configuration, and the controller and operator patterns behind self-healing.
Cloud Native Patterns
Designing Change-Tolerant Software
Cornelia Davis · 2019 · ISBN 9781617294297
What an application has to do differently when its infrastructure is expected to change underneath it: statelessness, configuration, service discovery, and request resilience.
Infrastructure as Code
Dynamic Systems for the Cloud Age
Kief Morris · 2016 · 2nd edition 2020
Treats infrastructure definitions as a codebase with the same expectations of testing, modularity, and pipelines. Good at explaining why a working script is not yet a platform.
Observability Engineering
Achieving Production Excellence
Charity Majors, Liz Fong-Jones, George Miranda · 2022 · ISBN 9781492076445
Why high-cardinality structured events beat dashboards of pre-aggregated metrics when the question is one nobody anticipated. Shapes what a platform should emit, not just what it should chart.
Delivery & DevOps · 7
Accelerate
The Science of Lean Software and DevOps
Nicole Forsgren, Jez Humble, Gene Kim · 2018 · ISBN 9781942788331
The research behind the four delivery metrics and the practices that move them. The book to cite when a leadership audience asks for evidence rather than opinion.
Continuous Delivery
Reliable Software Releases through Build, Test, and Deployment Automation
Jez Humble, David Farley · 2010 · ISBN 9780321601919
The deployment pipeline, defined. Older tooling in the examples, but the principles behind every modern CI/CD platform are laid out here first.
The Phoenix Project
A Novel About IT, DevOps, and Helping Your Business Win
Gene Kim, Kevin Behr, George Spafford · 2013 · ISBN 9780988262591
A novel, which is why non-technical leaders will finish it. It makes work in process, constraints, and unplanned work visible in a way a slide deck cannot.
The DevOps Handbook
How to Create World-Class Agility, Reliability, and Security in Technology Organizations
Gene Kim, Jez Humble, Patrick Debois, John Willis · 2016 · 2nd edition 2021
The practitioner companion to The Phoenix Project: the three ways, flow, feedback, and continual learning, with case studies for each practice.
Site Reliability Engineering
How Google Runs Production Systems
Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy · 2016 · ISBN 9781491929124
Where SLOs, error budgets, and blameless postmortems come from. Take the principles and leave the scale; most organizations need the chapters on toil and on-call, not the ones on planet-scale systems.
Docs for Developers
An Engineer's Field Guide to Technical Writing
Jared Bhatti, Zachary Sarah Corleissen, Jen Lambourne, David Nunez, Heidi Waterhouse · 2021 · ISBN 9781484272169
A practical process for documentation that engineers will actually follow, from reader research to maintenance. Useful for anyone who owns an API, a platform, or an architecture repository.
Chaos Engineering
System Resiliency in Practice
Casey Rosenthal, Nora Jones · 2020 · ISBN 9781492043867
Experiments against production as a discipline rather than a stunt, including how to get the organization to agree to them. Pairs well with the resilience patterns in this catalog.
Security · 4
Threat Modeling
Designing for Security
Adam Shostack · 2014 · ISBN 9781118809990
The standard text on asking what can go wrong before the design is finished, with STRIDE as the working checklist. Makes security review something an architect can run, not just request.
patterns DP-340
Security Engineering
A Guide to Building Dependable Distributed Systems
Ross Anderson · 2001 · 3rd edition 2020 · ISBN 9781119642787
The broadest single treatment of how systems fail against adversaries, from protocols to economics to human factors. Long; use it as a reference rather than a read-through.
Building Secure and Reliable Systems
Best Practices for Designing, Implementing, and Maintaining Systems
Heather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea, Adam Stubblefield · 2020 · ISBN 9781492083122
Makes the case that security and reliability are the same engineering problem and should share design principles, controls, and incident practice. The least-privilege and recovery chapters are the payoff.
Zero Trust Networks
Building Secure Systems in Untrusted Networks
Evan Gilman, Doug Barth · 2017 · 2nd edition 2024
What zero trust means once the marketing is removed: identity for every user, device, and workload, and policy enforced per request. Helps separate the architecture from the products sold under the label.
Organization & leadership · 6
Team Topologies
Organizing Business and Technology Teams for Fast Flow
Matthew Skelton, Manuel Pais · 2019 · ISBN 9781942788812
Four team types, three interaction modes, and the idea that team cognitive load is an architectural constraint. The vocabulary most platform and product org designs now use.
patterns DP-480
Thinking in Systems
A Primer
Donella H. Meadows · 2008 · ISBN 9781603580557
Stocks, flows, feedback loops, and leverage points, from outside the software field entirely. The mental model behind why most enterprise interventions push in the wrong place.
An Elegant Puzzle
Systems of Engineering Management
Will Larson · 2019 · ISBN 9781732265189
Engineering management treated as a set of systems to tune: team sizing, migrations, and how to run a reorganization without breaking delivery. The migrations chapter alone is worth it for an architect.
Staff Engineer
Leadership Beyond the Management Track
Will Larson · 2021
What senior technical roles actually do day to day, including the architect archetype, and how to get influence without authority. Honest about the parts of the job that are mostly writing and waiting.
The Manager's Path
A Guide for Tech Leaders Navigating Growth and Change
Camille Fournier · 2017 · ISBN 9781491973899
One chapter per rung from tech lead to CTO, which makes it the fastest way for an architect to understand what their engineering-management counterparts are being measured on.
Escaping the Build Trap
How Effective Product Management Creates Real Value
Melissa Perri · 2018 · ISBN 9781491973790
Why shipping features is not the same as delivering outcomes, and what a product operating model looks like. Architects who work with product organizations should know this argument cold.
Enterprise architecture · 5
Enterprise Architecture as Strategy
Creating a Foundation for Business Execution
Jeanne W. Ross, Peter Weill, David C. Robertson · 2006 · ISBN 9781591398394
Operating model first, then architecture: the diversification, coordination, replication, and unification quadrants still explain why two companies with the same ERP need different integration strategies.
The Software Architect Elevator
Redefining the Architect's Role in the Digital Enterprise
Gregor Hohpe · 2020 · ISBN 9781492077541
The architect as the person who rides between the executive penthouse and the engine room and carries meaning both ways. Full of reusable metaphors for explaining technical decisions to leadership.
patterns DP-500
A Seat at the Table
IT Leadership in the Age of Agility
Mark Schwartz · 2017
An argument from a former government CIO that IT leadership should own business outcomes rather than deliver requirements. Reframes governance, budgeting, and the vendor relationship accordingly.
Architecture Modernization
Socio-Technical Alignment of Software, Strategy, and Structure
Nick Tune · 2024
Ties domain boundaries, team boundaries, and investment strategy into one modernization approach, with collaborative techniques for discovering the domains instead of assuming them.
Just Enough Software Architecture
A Risk-Driven Approach
George Fairbanks · 2010 · ISBN 9780984618101
How much architecture work to do is a function of risk, not of process. The right antidote for both no-design agile and heavyweight documentation programs.
patterns DP-500