zgba 站群
How to write an effective software design document

How to write an effective software design document

A good design doc can save you years of development time. Writing a design doc forces you to think through important decisions before you waste time on the wrong implementation. It’s also the best way to coordinate design decisions among teammates and partner teams.

I’ve written design docs as a developer at Google, Microsoft, and within my own companies. The specifics vary, but the underlying principles remain the same. A design doc articulates the hard problems you’re solving and helps your teammates give you feedback.

Below, I share my approach to creating effective design docs and explain what belongs in a design doc and what does not.

The most common question I get about design docs is where to find a good one. I’ve never seen a public design doc that I consider high-quality. All of mine are hidden away at the companies that paid me to write them.

So, I wrote a design doc from scratch based on the principles I’m sharing here. It lays out the design for a real web app I’m building.

I created the design doc before writing any code, and I’m adhering to the design as I implement the app.

The design is more exhaustive than what I’d normally write for a solo hobby project, but this is roughly the length and depth of a design doc I’d create if I were coordinating work with other people on a professional project.

The more complex or risky the project, the more valuable it is to write a design doc.

Consider these questions:

If you answered “yes” to any of these questions, then it’s likely worth the effort to write a design doc. If you answered “yes” to two or more, a design doc will almost certainly be worth the effort.

A design doc can be a simple one-pager or a 50-page document that requires signoff from five different teams. You need to decide how much detail makes sense.

There’s no universal rule that says how long you should spend on a design doc just like there’s no rule that says how much to test your code. The right investment depends on your team’s goals, risks, deadlines, and culture. Sometimes, the right amount to invest in a design doc is zero.

If you specify every possible detail in a design doc, you’ve essentially written the implementation during the design phase. That would defeat the whole purpose of a design doc.

As a rule of thumb, you can ask a simple question to decide whether a decision belongs in your design doc: what’s the penalty for being wrong?

Not all design decisions are equally important. Some choices are more permanent than others.

For example, if you build a web application in C++ and realize 200k lines later that Ruby on Rails was the better choice, you’re stuck. A from-scratch rewrite would never work, and even if you manage to write new code in Rails, you’re still maintaining code in two wildly different languages.

Other design decisions are trivial. For example, if your app displays a list of 100 articles, should they all appear at once? Or should the user see 25 at a time and click “Load more” to see the next 25?

A “load more” button is not a design-level concern. If you pick one solution, and user feedback tells you you’re wrong, you can fix it in a few hours. You don’t need to detail your entire thought process in your design doc, and you definitely shouldn’t waste review cycles arguing about it.

Below, I’ve included common sections to include in your design docs. You generally don’t need every single section for every doc. Choose the subset that make sense for you.

The first thing your project needs is a title. It’s the way people will refer to your project in conversation, so aim for something short, distinctive, and evocative.

For example, if you were adding a caching layer between your application server and your database server, RecencyBank would be a good name. It’s easy to say and describes your project’s purpose. A bad name would be “Project Flying Silver Horse” because it’s verbose and nonsensical.

Boring but useful, metadata helps your reader understand the basic context of your doc:

The objective is a one-sentence explanation of your project’s purpose. It should appear on the first page of your doc in plain language that any stakeholder understands.

Improve application performance by adding a caching layer between the Trogdor web server and the Postgres database.

The background section explains the context and motivation for the project. It should answer these questions:

When we launched the Trogdor web app in 2023, pages typically loaded in 100ms or less. After three years, median page loads have ballooned to 600ms, which causes users to perceive our app as sluggish.

We investigated the slowdown and discovered that database lookups make up 80% of page load times. As our data store has grown larger, database lookups have gotten slower.

We also discovered that 95% of database lookups are for the same 3% of database rows. This pattern of usage benefits greatly from memory-backed caching. The cache would serve frequently-accessed data faster and reduce database load for all other queries.

Does your design doc make sense without outside context?

Imagine what you’d say to a teammate or partner team before they read your design doc.

Now, realize that some readers will see the doc before hearing any explanation from you, so whatever they need to understand should be on the first page of your doc.

If this project connects to other documents, make it easy for the reader to find them.

The goals section describes your high-level goals for this project. It should connect logically to the background section and explain what the world looks like after you’ve completed implementation.

Avoid setting goals in terms of implementation details. Your goals should communicate how the project benefits your users, your team, or your company.

While the goals define what’s within your project’s scope, the non-goals section delineates what’s out of scope.

Are there goals that readers might mistakenly assume are within scope for your project? If so, add them as explicit non-goals.

If your goal is something like “Add a ‘Share as URL’ button to charts,” the reader might not understand what that looks like in practice.

The scenarios section allows you to paint a picture for your reader of how your completed system works in the real world.

Scenario: Share a report via URL

Diagrams are tremendously valuable, though they might not seem that way.

As the design author, you intuitively understand how the pieces of your plan fit together. You can see the architecture in your head. Your reviewers do not have this mental picture, so the fastest way for them to see it is to draw them a picture.

Example diagram showing the architecture of a simple web application.

If you’re not sure what belongs in a diagram, think about these questions:

Choose a diagramming tool that’s flexible to editing. I’ve seen developers create a beautiful diagram on a whiteboard and photograph it for their design doc. The first draft looks amazing, but then they’re stuck with that diagram forever because they can’t edit the photo without recreating the whole thing from scratch.

Excalidraw, draw.io, and Google Drawings are popular diagramming tools that facilitate revisions. There are also languages like Mermaid, D2, and Graphviz that allow you to generate diagrams programmatically. I’ve had good experience using an LLM to create diagramming code for me. Remember to link to the source drawing or code so that your teammates have a way to reproduce the diagram as well.

The glossary defines terms that your readers might not recognize.

Think hard about the potential readers of your doc, especially newer team members and people outside of your immediate team. Will those readers understand the names of internal tools or systems your doc references?

When possible, use terms that your audience recognizes without having to refer to a glossary. Defining a term in a glossary is better than not defining it at all, but the best solution is to use recognizable terms or define them inline so that the reader doesn’t have to jump around your document.

If there are major constraints imposed on your design by your budget, clients, infrastructure, or dependencies, explain the constraints so the reader understands the context of your design choices.

Our servers are all RISC-V, so all code and dependencies must run on RISC-V architecture.

An SLO creates a measurable, objective metric for your system’s performance. You’ve probably heard of service level agreements (SLAs). SLAs are just SLOs plus financial penalties for falling short.

Within a company, you typically don’t financially penalize your co-workers for mistakes (although, wouldn’t that be kind of fun?). So, design docs define SLOs rather than SLAs.

Your manager might tell you that your app must be “performant on mobile,” but that’s vague. Your manager’s idea of “performant” might be <2ms of latency, and you don’t want to wait until code complete to find out. A well-defined SLO prevents ambiguity by expressing goals in concrete, objective terms.

The typical considerations for your SLO are:

Service level objectives

Once you nail down your SLOs (above), it’s time to think about how you’ll measure them in production.

The simplest way to verify that you’ve achieved your SLOs is to test manually. As your organization matures, you should automate monitoring to discover SLO failures immediately.

When defining your monitoring strategy, ask yourself these questions:

The following events will trigger a page to the on-call engineer:

The timeline section breaks your project into milestones. It specifies when you’ll deliver results to project stakeholders.

Choose milestones that create useful artifacts for stakeholders. For example, start with a UI that shows dummy data, and show that to clients first. If it turns out you misunderstood the client’s requirements, fake data lets you find out early rather than after you’ve already implemented all the plumbing to populate the UI with production data.

If you don’t know how to estimate project timelines, I highly recommend Joel Spolsky’s “Painless Software Schedules.” The article is 25 years old, but it remains my favorite software estimation strategy.

Your project exists to serve people or other software systems, so what do those interactions look like?

The Trogdor Server struct currently depends directly on a PostgresDB Go struct like this:

PostgresDB has the following exported methods:

We will create a Go interface type with the same API surface as PostgresDB:

We will implement a RecencyBank caching type that implements the same interface and wraps the backend PostgresDB struct. The RecencyBank implementation will cache reads from Postgres and forward requests to Postgres when they mutate state or depend on data not in the cache.

The only change to the Server implementation will be replacing the type of one member with the new interface:

The dependencies section should answer questions like:

It’s easy to overlook this section, but decisions about language, libraries, and infrastructure have a major impact on the complexity and long-term maintenance costs of your system.

Think deeply about which dependencies will be difficult to change after implementation. Don’t worry so much about the ones that swap out easily. It’s difficult to change languages or storage backends, but if you’re dissatisfied with the third-party service you use to send emails, you can replace it in an afternoon.

To build secure software, developers must integrate security into the full software lifecycle, starting at the design stage.

The security section should answer questions like:

Even if you think security threats are unlikely or irrelevant in your system, it’s still helpful to document your rationale. Your explanation might prompt reviewers to identify threats you overlooked.

RecencyBank must not accept direct requests from the public Internet, as it does not enforce any access control.

RecencyBank will run on a segregated network where it only accepts inbound requests from the Trogdor web server and can only make outbound requests to the Postgres server pool.

The privacy section is an opportunity to think through the sensitive data your system handles and what safeguards you’ll put in place to keep it secure. It should answer these questions:

RecencyBank contains the same sensitive user data as the Postgres database, so it inherits the privacy policy of our Postgres systems. In particular, engineers may only access RecencyBank systems in production with an associated bug number. Engineers must minimize the user data they access to only what is strictly required to investigate a bug.

If your system operates in a highly-regulated domain like finance or healthcare, the legal section helps you comply with relevant laws.

Even outside of regulated domains, think about whether your system could break the law if things go awry. Explain how you’ll steer clear

View original article