Perspective · 02

What the Data Model Settles

Why a schema is the one decision that every later decision inherits.

The data model is chosen early, usually under time pressure, often by whoever is closest to the first feature. It is treated as an implementation detail because at the moment it is made it looks like one: a handful of tables, obvious relationships, nothing that appears to foreclose anything. It is the most consequential decision in the system, and it is taken when the least is known.

What follows inherits it. Queries are shaped by it. Interfaces expose it. Reports assume it. Integrations are written against it. By the time the domain is understood well enough for the model to be designed properly, the model has acquired dependents that make change expensive in a way that is difficult to explain to anyone who has not attempted one.

The first use case is not the shape

A model derived from the first use case describes that use case accurately and the domain only incidentally. The second use case is where the difference appears. It arrives carrying a requirement that is obvious in the domain and awkward in the schema: an entity that turns out to have two parents, a relationship that was one-to-many and is now many-to-many, a field that meant one thing and now means two.

The response, under pressure, is to accommodate rather than remodel. A nullable column. A type flag. A second table duplicating most of the first. Each accommodation is locally reasonable and cheaper than the alternative on the day it is made. Their accumulation is what people are describing when they say a system has become difficult to change.

The alternative is not to design for imagined futures, which produces its own expensive generality. It is to model the domain rather than the feature — to ask what the things are, independently of what the first screen needs to display.

Identity, and what it costs to get wrong

The choices hardest to reverse concern identity: what constitutes a distinct thing, and what makes two records the same thing.

Identity decisions propagate outward. They determine what can be counted, what can be merged, and what happens when the world disagrees with the model — when the same organisation appears twice under different names, when a record must exist before it has the attributes that identify it, when two systems that have to reconcile hold different notions of sameness.

Any schema answers these questions, explicitly or otherwise. A model that gives an entity a surrogate key has said one thing. A model that keys on an attribute has said another, and has also said that the attribute will not change. Attributes chosen as keys have a way of changing.

What a migration cannot recover

Schema change is discussed as though it were a matter of care and downtime. Structural change is. Semantic change is not.

If a column has meant two things at different times, no migration recovers which rows meant which. If a relationship was maintained by convention rather than by constraint, the violations already in the data are indistinguishable from the intended cases. If a value was overwritten rather than versioned, the prior state is simply gone, and no amount of care applied later retrieves it.

This is the argument for constraints, and it is not an aesthetic one. A constraint is a claim about the domain, enforced at the only point where enforcement is reliable. A claim not declared in the schema is one maintained by discipline across every path that writes — including the paths written later, by people who did not know the claim existed.

The model as a record of decisions

A well-formed schema documents the decisions an organisation has made about its own domain: what is a distinct thing, what may be absent, what may not, what has to be true simultaneously. Read carefully, it is the most honest description of how the organisation understands its own work, because it is the description the software is actually obeying.

This is why the model is the first thing worth reading and the first thing worth designing. It is also why a schema that is difficult to read is a symptom rather than a matter of style. A model nobody can explain is one where the decisions were made incidentally, and a system whose foundational decisions were incidental will surprise the people operating it — usually at the point where surprise is least affordable.

Settled first, not settled forever

Settling the model first does not mean freezing it. It means taking the decisions deliberately, while they are still cheap, and recording the reasoning — because the reasoning is what a later change needs and what is otherwise lost with the people who held it.

The practical form of this is unglamorous. Name the entities before the screens. Declare the constraints that are actually true. Version what will be asked about later. Treat a required migration as a design event rather than a maintenance task. None of it is difficult. It is simply earlier than most schedules want it to be, and it competes for attention with the demonstration that everyone can see.