Entities give us the vocabulary of a system. Relationships give it grammar.
A model composed only of entities is a list of nouns with no verbs. The interesting behavior of any real system — the behavior that is hard to change without breaking something — lives almost entirely in how entities connect to one another. Getting entities wrong makes a system confusing. Getting relationships wrong makes it dangerous: data diverges silently, and the system acquires behavior that nobody designed and nobody can fully explain.
Relationships deserve the same deliberate treatment we gave entities in Entities — The Things That Exist. That means making them explicit and encoding the rules that govern them before those rules get buried in application code.
A relationship encodes how two entities are linked and what that link means. Three examples from our running systems:
Robot is assigned to a Delivery_TaskDocument carries a TagWorld_Object interacts with another World_ObjectIn each case, the link is not decorative. It defines which operations are permitted and what must remain consistent across the system. An assignment, for example, implies that the robot is unavailable for other tasks, and an interaction between world objects means that the rules governing both must be consulted whenever either one changes state.
When these implications are left implicit — encoded nowhere in the model, enforced nowhere in the code — they do not disappear. They simply become assumptions that every developer must independently discover and manually respect. That is the source of most subtle, long-lived bugs in large systems.
To make a relationship fully explicit, five questions must be answered.
Cardinality. How many instances of each entity can participate? A robot may be assigned to at most one active delivery task at a time — this is a one-to-one constraint. A document may carry many tags, and a tag may apply to many documents — this is many-to-many. Getting cardinality wrong produces either artificial restrictions or data models that permit states the domain forbids.
Direction. Is the relationship symmetric or does it flow one way? A delivery task references an origin and a destination, but a location does not inherently reference the tasks that pass through it. A document link between A and B may or may not imply a link from B to A, depending on the link type. Direction determines which entity “owns” the relationship and which follows from it.
Ownership. Which entity is responsible for maintaining the integrity of the link? In some relationships, one entity is the parent and the other is a dependent — the parent’s existence is a precondition for the child’s. In others, the relationship is mediated by a shared index or junction structure, and neither entity fully owns it. Leaving ownership undefined means that two different parts of the system may each assume the other is maintaining the link, and neither does.
Lifecycle coupling. What happens to a relationship when one of its participants is removed? Three outcomes are possible: the related entity is deleted with it (cascade), the related entity is left without a valid reference (orphan), or the relationship is preserved with a historical marker that makes the absence explicit. None of them is automatically correct, but each must be decided rather than discovered.
Constraint rules. What must always be true across the relationship? A link must reference entities that actually exist, and certain link or relationship types may forbid self-reference or cycles. These constraints are invariants of the relationship itself, and they belong in the model for the same reason that entity invariants do.
Applying these dimensions to the running systems shows how differently relationships can look while answering the same questions.
| System | Key relationships | What the model must specify |
|---|---|---|
| Delivery network | Location to location (paths); robot to task (assignment); task to origin and destination | Each path has a direction and a status. A robot holds at most one active task, and the assignment is released when the task completes or fails. A task is valid only if both of its locations exist. |
| Knowledge engine | Document to tag; document to document (links), including links the system infers | The document–tag association is many-to-many and can carry confidence or provenance. Each link is typed, and the type determines how traversal treats it. Inferred links carry a confidence score that distinguishes them from links a user asserted. |
| Virtual world | Object to region (containment); interaction rule to a pair of object types; event to the transition it caused | Containment decides which interaction rules apply. Rules relate types rather than individual objects, linking the type system to the entity system. Event links record cause and effect, so a transition can be traced to the event that triggered it. |
None of these relationships is an incidental field on an entity. Each has its own constraints and lifecycle, and the model has to state them.
Sometimes a relationship is simple enough to be represented as a reference field on one of the participating entities. A delivery task carries an identifier for its assigned robot. A document carries a list of tag identifiers. These representations work when the relationship has no properties of its own and when only one query direction matters.
When a relationship needs to carry metadata — a timestamp, a confidence score, a link type, a source — it becomes a thing in its own right and deserves to be modeled as an entity. Consider the requirement:
“The knowledge engine should connect related notes.”
The first candidate model is a direct many-to-many connection: each document holds a list of related document identifiers. This works until we need to know when a connection was made or what kind of connection it is. At that point the flat field becomes inadequate, because the relationship has acquired properties that cannot live on either endpoint.
The stronger model names the relationship explicitly. Start with the entities: Document, Tag, Link. A Link connects two documents with a typed edge. Constraints follow: both endpoints must exist, and a link’s type must come from a controlled set. Query patterns follow from there, including backlinks and two-hop traversals that surface indirectly related material.
The discipline is to model for real query patterns, not for theoretical elegance. A model that supports forward lookup efficiently but makes reverse lookup expensive or impossible is a model that was designed without asking how the data would actually be used.
The following sketch expresses the Doc_Link relationship entity in Nex:
class Doc_Link
feature
from_id: String
to_id: String
link_type: String
is_structurally_valid(): Boolean do
result := from_id /= "" and to_id /= ""
and link_type /= ""
end
create
make(from_id, to_id, link_type: String) do
this.from_id := from_id
this.to_id := to_id
this.link_type := link_type
end
invariant
endpoints_present: from_id /= "" and to_id /= ""
non_self_reference: from_id /= to_id
link_type_present: link_type /= ""
end
This is minimal by design. It captures the invariants every link must satisfy and exposes a validation operation whose contract is explicit. The underlying document objects are referenced by identifier rather than by direct containment, which means the link’s validity depends on those documents existing in the broader model, not just in this class. That dependency belongs in the integration-level invariants we write when assembling the full model.
The same structure generalizes: Path in the delivery network is a relationship entity between locations; typed interaction edges in the virtual world are relationship entities between object types.
Encoding relationships as free-text fields. When entity references are stored as untyped strings, such as names or other human-readable identifiers, the model has no way to enforce that the referenced entity actually exists or that two references to “the same thing” are in fact consistent. Joins become unreliable and traversals become impossible. The recovery is to model relationships as typed links with constrained endpoints.
Ignoring reverse queries. A relationship that is easy to traverse in one direction and expensive or impossible to traverse in the other was designed for only half of its intended use. Forward lookup and reverse lookup are both first-class access patterns, and both should be considered when the relationship is modeled. If the reverse query is expensive under the current model, that is information about the model, not about the query.
Semantic drift in link types. When the same link type accumulates subtly different meanings across different parts of the system, the relationship becomes uninterpretable: a traversal that was valid under the original meaning may be invalid under the acquired one, and there is no record of when or why the meaning changed. The recovery is to define a controlled taxonomy of link types at the model level and to enforce membership in that taxonomy when links are created.
Hidden lifecycle rules. When an entity is removed and the links that referenced it are not updated, the system accumulates broken references that will produce failures at some unpredictable future point. The lifecycle policy — cascade, orphan, or preserve — must be defined explicitly at the model level and verified by tests that exercise entity removal, not just entity creation.
Quick Exercise
Choose one of the three running systems and construct a relationship matrix. For each relationship you identify, record: the two entity types it connects, the relationship type and cardinality, and one constraint rule that must always hold.
Then identify one reverse query that your model must support — a query that traverses the relationship in the direction opposite to how you first defined it. If that reverse query is expensive or ambiguous under your current model, the model needs refinement before implementation begins.
Takeaways
The next chapter, Designing a Good Data Model, brings entities and relationships together into a complete data model, and examines the tradeoffs that arise when a model must serve multiple competing concerns simultaneously.