Errors and measurement · 14 min read
Why no resolver gets every link right
One change to a hypothetical pipeline breaks three links nobody touched. What makes a link correct, why are some incorrect and missing links certain, and when can a resolver be left to add links on its own?
Roar Elias Georgsen, 27 September 2026
Part 3 of 7 in Automating traceability.
The story so far
Part 2 wrote a resolution task down as a card, from who uses the links to the version of each side a link is resolved against. However carefully that card is filled in, the resolver that carries it out will propose some incorrect links and miss some correct ones. This part asks what a correct link is, and what can fairly be expected of a resolver.
Suppose the hypothetical pipeline from the first two parts has been built and running happily for a while. The task from part 2 has linked the parse stage, PIPE-S2, to the OpenTofu resource module.pipeline.aws_ecs_service.stage["parse"], and the link is correct, since that resource really does deploy parse. Then the team that built the pipeline splits parsing in two. Tokenising moves into a new lex stage, and the list of stages in the configuration gains a fourth name. In the same change, the service that reported to OpenTelemetry as query-parser starts reporting as query-frontend. Nobody on either team touches the systems model, so PIPE-S2 is still one part.
Nobody edits a link either, and three of them change meaning anyway. The OpenTofu link still resolves, because the parse instance still has its address, but that instance now does half of what the part used to do. In telemetry, the old link points at a service that has stopped sending anything. The new query-frontend traces resolve to nothing, since nothing in the service name says parse. Meanwhile the Checkly check parse-sustained-qps goes on measuring throughput, now through two services where the systems model says one.
That’s the decay from part 1, and it cuts both ways. Mäder and Gotel describe links that, left unmaintained, “get lost or represent false dependencies”, so the same neglect produces missing links and incorrect ones. Tools that manage trace matrices mark a link as suspect when either end changes. Cleland-Huang and colleagues’ review of the field notes that “it is not uncommon to see an industrial trace matrix populated with a high percentage of suspect links”.
My own records managed it with just two documents. Two of the demo’s decision records gave the same detail about how the compose input sets up subscriptions, and they agreed with each other. When I corrected the first on 27 August, they stopped agreeing. The second went on saying the old thing for 28 days, until 24 September, and nobody touched it in between. I’d like to say I noticed sooner. A link that was correct once only tells you it was correct then.
What makes a link right
“Right” can mean several things when it’s said of a link. Two of them matter most here, and the rest of this series has to tell them apart. A link is correct when the relationship it claims holds in the built system as it stands. It can also be, or fail to be, the link a task needs. Part 1’s check parse-sustained-qps measures throughput through parse, so a link from it to PIPE-S2 would be correct. A coverage report still needs the check linked to the requirement it verifies, PIPE-R1.2, and gets nothing from the link to the part.
Key term Correct, incorrect and missing links
A correct link claims a relationship that holds in the built system as it stands. An incorrect link claims one that doesn’t, and counts as a false positive. A missing link is a correct link nobody has made, and counts as a false negative. A correct link can still be one the task doesn’t need, and the task card from part 2 says which kind it does.
Moving tokenising out of parse into its own lex stage and renaming the parse service produced one of each, without anyone touching a link. The OpenTofu link is still correct, although the instance it points at now does half of what the part used to do. The link to query-parser has become incorrect, since that service no longer sends anything. And the query-frontend traces stand for a missing link, correct and not yet made.
How many are still correct?
After a change like that one, the obvious question is how many of the pipeline’s links are still correct, and nobody is going to check every one of them by hand. A fixed benchmark won’t settle it either. One built from last year’s system decays along with the system, and one built from a few dozen links may never have looked much like the real thing.
Sampling does the job. Tepping proposed a way in 1968, as Binette and Steorts describe it. Group the pairs into classes by which of their fields agree, check a sample from each class by hand, and estimate the error rate class by class. Then split further the classes that add most to the expected cost. The estimate is only as good as the sample, and it has to be taken again as both sides change. Part 4 is about measuring properly, and it comes back to this.
Names are written for people
A good enough resolver might seem able to resolve every name correctly, if only someone built it carefully. My demo suggests otherwise. Here are eight of the names in it that contain the word router:
Router, a part definition in the demo’s systems model, which works like a typerouter, the part in the same systems model that is declared from it- ROUTER, the label that definition’s documentation gives the router when requirements are allocated to it
sysml-federation-router, the service name in the router’s telemetry configurationrouterVersion, an attribute holding the router’s version number, 0.343.1router-version, the logical ID of a Checkly check- “router: the model version”, the same check’s display name
CHK_RouterVersion, the verification case in the systems model that describes that check
The first four are the router, in four different spellings, and the other four aren’t. What makes this a trap is the pair in the middle. routerVersion holds the version of the router, while the check called router-version asks the demo for the version of the systems model it serves, through the router. A name strategy sees “router” in all eight and “version” in four. I chose every one of these names myself, which rather spoils the option of blaming anyone else.
Even a careful person can go round in circles on one of them. The check’s own file calls it one of “the seven API checks on the router”, and it sits in the router’s group. The systems model’s verification case for the same check gives its subject as the whole demo. It says the check verifies SR-43, which asks that one response carry a requirement’s text, verdict and document number from a schema all three services contribute to. My two records disagree about what the check is about, and I wrote both.
Names in a built system are written for the people who work on it, and those people read them with context a resolver doesn’t have. Text-based recovery has leaned from the start on the premise, in Antoniol and colleagues’ words, “that programmers use meaningful names for program items”. Mostly they do, and the names mean something to whoever chose them. The 2014 review found that the gains from those methods “seem to have plateaued”, mostly because of term mismatches between the documents being traced.
Resolvers still differ a lot. One that reads the stage tag on the OpenTofu resource still links parse correctly after tokenising moves into lex, because the tag still says parse. One that compares service names is stuck the moment query-frontend appears. The difference is which clue survived this particular change, and the next change may favour the other resolver.
Sometimes the fault isn’t in the resolver at all. Once tokenising had its own stage, the new lex instance had nothing to resolve to, because the systems model has no lex. A resolver that trusts the systems model completely drops that instance without comment, and yet it’s the most useful thing the run turned up. The systems model is out of date, and the unresolved object is the evidence. I’d have the resolver report it as a suspected mismatch between the systems model and what was built.
Researchers in software architecture met this in the 1990s. Murphy and Notkin’s reflexion models compare a high-level diagram of a software architecture with the code. An engineer writes a map from the code to the diagram’s boxes by hand, often with regular expressions over file names, and a tool then reports where the two agree and where they don’t. In their case study, the first comparison found 15 convergences, 83 divergences and four absences. Where the code had an interaction the diagram lacked, the engineer updated the diagram. My own systems model has been the stale side too. When I read each of the demo’s 36 drawings and screenshots against the views in its systems model on 12 September, I found three things in them that the systems model didn’t yet hold. So an object that resolves to nothing is a report on the systems model as much as on the resolver, and whichever way of working you choose, that report should reach a person.
Fixing whatever gets reported
The morning after the split of the parse stage, its service is reporting as query-frontend and tokenising runs in its own lex stage. Someone notices that the query-frontend traces link to nothing and asks for them to be linked to the parse stage. It’s a fair request, and there are two quick fixes.
The first is a rule saying that query-frontend is PIPE-S2. It fixes exactly the reported miss and nothing else, so on today’s links it costs nothing. It’s also a key written by hand, the kind part 1 found holding my demo together in places, and it decays like any other hand-written link. The next time someone renames the service, the rule finds nothing and the miss comes back. A resolver patched one complaint at a time ends up as a list of hand-written links with a scoring function bolted on.
The second fix lowers the threshold on the name strategy until query-frontend scores high enough. Its cost is harder to see, and it needs two measures.
Key term Precision and recall
Two measures of the links a resolver proposes, taken against the links that ought to exist. Precision is the share of proposed links that are correct, so every false positive lowers it. Recall is the share of the links that ought to exist that the resolver found, so every false negative lowers it.
Say the telemetry task on a bigger version of the pipeline ought to produce 40 links, and its name strategy scores each candidate between 0 and 1. The numbers below are an example, chosen to show the shape of the trade.
| Threshold | Links proposed | Of which correct | Precision | Recall |
|---|---|---|---|---|
| 0.8 | 25 | 23 | 0.92 | 0.58 |
| 0.5 | 60 | 36 | 0.60 | 0.90 |
Lowering the threshold from 0.8 to 0.5 to catch one reported miss brings in 35 more links. Of those, 13 are correct and 22 incorrect, and nobody will report the 22, because the person who asked for the fix was looking at something else. It works the other way too. Raise the threshold to get rid of a reported incorrect link, and correct links go with it. Fixes like these pile up, each tuned to a mistake somebody happened to notice, and each needing testing against everything the resolver already does. The links nobody looked at pay for them. The thing to aim for is the best balance across all the links, and which balance is best depends on who pays for each kind of mistake.
What an incorrect link costs
Suppose PIPE-R1.2, the throughput the parse stage must sustain, ended up linked to a latency check called parse-p95-latency. A coverage report would show the requirement verified, by a check that never measures throughput. A missing link would have shown it unverified, which is a gap somebody would ask about. I haven’t found a source that says this in so many words, so take it as my argument. It fits what Mäder and colleagues report of the traceability sent to regulators, which “is often weak, casting doubt rather than confidence”.
While a person is reviewing candidates, it’s the other way round. An incorrect candidate takes a moment to reject, and part 1 quoted Dekhtyar and Hayes on how much faster that is than finding an omission. They go on to call recall in tracing “significantly more important than precision”. Rodriguez and colleagues add that “many safety-critical domains require near-perfect recall when using NLP techniques to automatically generate links, which can not be consistently achieved”. In practice that pushes safety work towards suggesting links for a person to confirm.
Record linkage has a formal version of the argument. Tepping’s decision rule, in Binette and Steorts’s account, gives each action one cost when the pair is a true link and another when it isn’t, and picks the action with the lowest expected cost. Unequal costs move the thresholds. So my position is part 1’s, with a reason attached. When a person will see every candidate, favour recall. When links go straight into anything a coverage report or a safety case reads, favour precision, because that’s where an incorrect link hides.
It follows that a resolver adding links on its own should be free to say it doesn’t know. Record linkage has that already, in part 2’s middle band of pairs held back for a person, and machine learning has called it a reject option since Chow’s work in 1970. The query-frontend traces above show why a missing link is the safer mistake. Having nothing to link them to, the resolver linked nothing, and somebody asked about the gap the next morning.
The person seeing every candidate isn’t a perfect filter, either. Part 1 reported that analysts improve poor candidate sets and decide worse on good ones, and Cleland-Huang and colleagues add that human feedback on trace links “is incorrect approximately 25% of the time”. Continuous resolution brings a trap of its own. A resolver that runs on every change will want to show a reviewer only what changed since the last run, which spares the reviewer every link that hasn’t moved. The same review reports that showing analysts only new or changed links “negatively impacts the quality” of the links. A resolver that works this way should at least say when each link was last seen by a person in its full context.
Letting a resolver add links on its own
Now say someone adds lex to the systems model, and the resolver is sure the new OpenTofu instance belongs to it. I’d let it add that link without asking anyone, under a few conditions. A blanket rule against it compares automatic links with a perfect hand-made set that doesn’t exist. Rath and colleagues found that on average about 40% of commits carried no issue key, in projects that largely followed the practice of tagging every one. The links already in my demo’s systems model came from me, by hand, and this part has already caught my records contradicting each other twice.
The first condition is that every link says where it came from. Removal depends on it, since you can’t clean up after a bad run you can’t find. Guo and colleagues argue that maintenance has to handle “a mixture of automatically generated and manually generated trace links and leave the manually created ones untouched”. Part 5 comes back to this under the name provenance.
The second is that removing an incorrect link is cheap, and that the removal is remembered. Otherwise a resolver that runs on every change adds the same incorrect link back on its next run, and the person who removed it soon stops bothering. A rejected link is a result, and it needs recording as carefully as an accepted one. The third is that the automatic mode favours precision, for the reasons above. The fourth is the rule the trace-matrix tools above already follow. A link goes suspect the moment either of its ends changes, and stays suspect until something checks it again.
Taken together, those conditions make a link more than yes or no. It has a state, such as proposed, accepted, rejected, added automatically or suspect, and a score that says how sure the resolver was. A link with a state and a score can be trusted as far as its history supports, which is further than a bare cell in a matrix.
More links aren’t better
A resolver judged by how many links it adds will add too many. Heindl and Biffl found that tracing requirements by their value took around 35% of the effort of tracing all of them in full. The risky and volatile requirements were the ones that warranted more detail. At a 2015 Dagstuhl seminar on traceability, Cleland-Huang argued that the traceability certifiers prescribe “tends to be overly extensive”. For the parse stage, the link worth maintaining runs from PIPE-R1.2 to the check that measures throughput. A link from PIPE-S2 to every log line that mentions parsing adds nothing a search couldn’t.
Part 4 is about measuring a resolver against a set of correct links that changes along with the system.
Previous: The parts of a resolution task · Index: Automating traceability · Next: Measuring a resolver