Running a resolver · 15 min read
What makes a resolver worth running?
A resolver that measures well can still be a poor choice to run. What else decides it, from where a link came from to who may read the systems model?
Roar Elias Georgsen, 27 September 2026
Part 5 of 7 in Automating traceability.
The story so far
Part 4 measured a resolver against a gold set that moves with the system, with an F-score chosen for the way the resolver is put to work. A resolver that measures well can still be a poor choice to run. This part asks what else decides it.
Suppose the hypothetical pipeline from the earlier parts had been built, and a resolver had added the link from PIPE-S2 to module.pipeline.aws_ecs_service.stage["parse"] on its own. Six weeks later, an engineer decides the link is incorrect. Before removing it, they’ll want to know a few things. Which run added it, and which version of the systems model and which OpenTofu plan did that run read? Which rule proposed the link, and did a person ever confirm it? If all they have is a row in a table, they can delete the row. They still won’t know whether the next run will put it back, or whether a hundred other links came from the same fault.
Where a link came from, and where it lives
Key term Provenance
The record of how a piece of data came to be: what produced it, from which inputs, when, and who has vouched for it since. For a trace link, it’s what lets a person judge how far to trust the link, and undo a whole run’s links if the run was faulty.
The W3C’s PROV ontology already has words for all of that, and a trace link fits them without strain. Here’s what a record for the link from PIPE-S2 could look like:
@prefix prov: <http://www.w3.org/ns/prov#> .
@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .
@prefix : <https://example.org/trace#> .
:link-7f3a a :TraceLink, prov:Entity ;
:from "PIPE-S2" ;
:to "module.pipeline.aws_ecs_service.stage[\"parse\"]" ;
:state "added" ;
prov:wasGeneratedBy :run-1412 ;
prov:generatedAtTime "2026-10-02T14:12:09Z"^^xsd:dateTime .
:run-1412 a prov:Activity ;
prov:used :systems-model-9c1e2d4, :tofu-plan-118 ;
prov:qualifiedAssociation [
prov:agent :resolver-0.1 ;
prov:hadPlan :stage-name-rule-v3 ] .
:resolver-0.1 a prov:SoftwareAgent .
:stage-name-rule-v3 a prov:Plan .
In words, the run used a commit of the systems model and an OpenTofu plan, the resolver was a piece of software, and the rule it followed was, in PROV’s terms, a plan. That’s two different kinds of plan in one record, and I’m blaming the standards for it.
A few more statements complete the story over time. When a person confirms the link from PIPE-S2, that’s an activity of its own, associated with them. When someone removes it, that’s an invalidation, which PROV-DM defines as “the start of the destruction, cessation, or expiry of an existing entity by an activity”. The link stays in the record, marked invalid by a named activity, so the next run can see that someone rejected it. That’s the remembered rejection from part 3, in a standard vocabulary. Regulated work asks for the same thing in its own words. The FDA’s rule on electronic records, 21 CFR 11.10(e), requires time-stamped audit trails, and says that “record changes shall not obscure previously recorded information”.
A suggested link needs its evidence on show at the moment someone decides. A link added automatically needs its run on record, so that a bad run can be found and undone in one go. Somebody still answers for a link nobody confirmed, and I’d say it’s whoever decided that links of its kind could go in without a person looking. That decision is the only human one in the link’s history.
The record also needs a home. In my demo, each subgraph owns its own fields, and the router joins them on keys from the systems model. Resolved links fit the same pattern. A trace subgraph would add link fields to the graph’s types for elements of the systems model, keyed on their short names. Each link would carry its provenance, and the rest of the graph wouldn’t need to know that a resolver made them. That’s where part 1’s two ends meet. Entity resolution works out the keys, and federation joins on them.
Ownership comes with it. By default, Federation 2 doesn’t let two subgraphs resolve the same field. Apollo’s documentation says a single object field “can’t be defined or resolved by more than one subgraph schema”. Cosmo’s documentation treats sharing one without the @shareable directive as a composition error. If the trace subgraph owns the link fields, no other service can publish its own version of them without the graph refusing to compose.
The join rests on the short names, and my own records name the weak spot. The decision that made short names the keys says that “renaming a short name is a breaking change for every service that stores it, and the demo does not handle it”. A trace subgraph would store a great many short names, so renaming one is the first thing it would have to handle, and the demo doesn’t handle it yet.
Explaining a link
The capacity service in my demo explains every verdict with one of a few fixed templates, such as <quantity> <value> against <limit>, limited by <cut>. Nobody has to wonder where a reason came from, because there’s only one place it can have come from. A rule that links a test to SR-22 because its name starts with TestSR22_ can do the same, in the same words on every run. Anyone who disagrees can see which rule to change.
A learnt score gives less away. Say a learnt strategy links the Checkly browser spec refusals.multistep.spec.ts, whose file name carries no key, to a user story, with an example score of 0.81 against a threshold of 0.75. It lists the words that weighed most. That’s something, but both numbers belong to one version of the strategy. Sculley and colleagues point out that when a machine learning model updates on new data, “the old manually set threshold may be invalid”. The explanation a reviewer read last month may not be the one the resolver would give today.
A good explanation also tells the engineer what to change. “No link, because this resource has no stage tag” is a problem they can fix in a minute, and the next run links it correctly. A score of 0.62 tells them nothing they can act on.
Work that ends in certification wants more again. The certification scenario in Cleland-Huang and colleagues’ review of the field asks for “supporting rationales explaining how these artifacts mitigate the hazard”. An assessor will want the reason for a link as much as the link. Language models change the picture, because they write fluent explanations of their decisions. In one study, Rodriguez, Dearstyne and Cleland-Huang found that a language model “could provide an in-depth analysis of its decision”. They also said plainly that whether such explanations “are accurate reflections of the reasoning behind the model’s decision is beyond the scope of this paper”. Parts 6 and 7 pick that question up.
A reviewer confirming candidates needs a reason for each one at the moment of deciding, or they’re guessing. For a link added automatically, the reason matters later, when someone disputes it, and then it has to be the reason the resolver actually had. That’s one more thing provenance should keep.
Running on every change
My demo’s trace test reads the demo’s whole systems model and the repository, and its package finishes in a few hundredths of a second. Every pull request runs it afresh, with Go’s test cache out of the way. That’s the bar I’d hold a resolver to, since part 1 argued that resolution has to rerun whenever either side changes, the way a test suite does.
Most of that speed comes from reading keys. Part 2 noted that comparing everything with everything grows with the product of the two sides. Take an example system with 500 elements in its systems model and 20,000 built objects. An all-pairs scorer faces 10 million pairs, while a key rule needs 20,000 lookups in an index of 500 keys. Blocking, as Papadakis and colleagues put it, “trades slightly lower effectiveness for significantly higher efficiency”, and in a pipeline that runs on every commit, that trade usually pays. Cost adds up the same way. Compute and memory are paid on every run, retraining on every retrain, and the most accurate resolver is no use if the team won’t pay for it a hundred times a week. I’d keep an eye on those numbers over time, too, since they grow with the system.
Each way of working sets its own deadline. Suggesting a key while an engineer writes an OpenTofu resource has to happen while they’re still typing, in well under a second. Adding links in bulk after a merge can take minutes, as long as it finishes before the next merge. Recomputing only what changed is cheap and safe for the machine, but part 3 found that reviewing only what changed costs quality, so I’d still show a reviewer each link in its context.
A resolver can also break with nothing changed on either end of a link. Say a strategy tells the parse stage’s write path from its reads by the http.method attribute on its spans. One day the instrumentation switches to OpenTelemetry’s stable names, http.method becomes http.request.method, and the strategy finds nothing. No test fails, and nobody notices. That exact rename is in OpenTelemetry’s HTTP migration guide, which has instrumentations emit the old names, the new ones or both, depending on an environment variable. OpenTelemetry’s telemetry schemas exist for this. They define the transformations from one version to the next, and a resolver that normalises through them first would keep working.
OpenTofu is kinder. tofu providers schema -json versions each provider’s schema, and the dependency lock file records the exact provider version chosen, so a resolver can at least tell when the ground has moved. Part 3’s change to the pipeline shows the difference. When tokenising moved out of parse into its own lex stage, OpenTofu could record the move in a moved block, while the telemetry kept no record that query-parser became query-frontend.
New objects are the easier half of change. A rule or a key index picks up a new resource the moment it appears. A learnt model may need retraining before it knows about something new, although some can update as they go. For a resolver meant to run for years, that’s one of the first things I’d ask about.
My demo has a silent pass of the same kind on record. The decision that turns the router’s telemetry off depends on three environment variables the vendor doesn’t document, and notes that a release renaming one “would announce it nowhere. Nor would anything here notice.” A resolver built on names it doesn’t own should expect that, and check the versions of the formats it reads as carefully as the links it makes.
Who may read the systems model
My demo’s image makes no outbound connection on its own. I ran it with its network removed to check. Its optional check session is a different matter, and deliberately so. Checkly’s hosted runners reach the demo through a public tunnel and read the example pipeline’s systems model, which is fine for a pipeline that doesn’t exist. A confidential systems model would need Checkly’s private location, which puts the runners beside the demo, and I’ve never run it because it needs a paid plan.
A resolver faces the same choice, and in some industries the rules make it for you. In US defence work, a systems model is likely to count as technical data. The export rules in 22 CFR 120.33 define that as information “required for the design, development, production, … or modification of defense articles”. The rules exempt some data sent with end-to-end encryption, under conditions that include, in 120.54, that “the means of decryption are not provided to any third party”. A hosted resolver has to read the systems model in clear, so as far as I can tell it can’t use that exemption. Whether it counts as an export then turns on where it runs and who can reach it. I’m not a lawyer, and this is my reading of the text. When in doubt, I’d run the resolver where the systems model lives.
Open and reproducible
An auditor asks why PIPE-S2 is linked to the parse service. The best answer is to rerun the resolver on the same inputs and get the same link. That’s the trace-link version of a reproducible build, one where “given the same source code, build environment and build instructions, any party can recreate bit-by-bit identical copies of all specified artifacts”. For a resolver, the same commit of the systems model, the same OpenTofu plan and the same resolver version should give the same links, which rules out any randomness that isn’t pinned down.
Being able to read the resolver’s code makes that much easier to check, and open code brings more. Other people test it on their own systems, fix what breaks and improve it, and you’re free to change it yourself when your needs differ. The Open Source Definition adds a clause that matters here. An open licence “must not restrict anyone from making use of the program in a specific field of endeavor”, and the fields in question include defence and medical devices. My decision records have an example of the opposite. A SysML tool I considered for the demo was closed source and licence-gated, “with air-gapped use needing a Business plan”, and air-gapped is exactly how defence and medical work tends to run. Open code needs documentation and tests beside it, or nobody else can check what it does, and checking is most of the point of being able to read it.
Choosing one
Here’s how I’d decide for the parse task, with candidates and scores chosen as an example. Before looking at any results, I’d add a line to the task card from part 2:
Gate: F0.5 >= 0.8, under 60 s in CI, on premises
A hosted service scores best, but the systems model can’t leave the building, so it fails the gate. A learnt model scores well too, but retraining it takes an hour after every merge, so it fails as well. That leaves two, the rule cascade from part 2 at an F0.5 of 0.84, and the same cascade with a learnt model added at 0.85.
Choosing between those two is a judgement, and SysML v2 has a place to write a judgement down. Its domain library defines a trade study as an analysis whose subject is a set of alternatives. An evaluation function scores each one, and an objective says whether the best score is the highest or the lowest. The gate becomes a requirement on the resolver, and the choice becomes an analysis in the systems model:
requirement def Gate {
subject r : Resolver;
require constraint { r.f05 >= 0.8 }
require constraint { r.ciSeconds < 60.0 }
require constraint { r.onPremises }
}
analysis choice : TradeStudy {
subject :>> studyAlternatives : Resolver[2] =
(cascade, cascadePlusLearnt);
calc :>> evaluationFunction {
in ref :>> alternative : Resolver;
return :>> result : Real =
alternative.f05 - 0.02 * alternative.strategies;
}
objective :>> tradeStudyObjective : MaximizeObjective;
}
Both SysML v2 reference tools accept it, once the resolver’s definition and the two candidates are added. The evaluation function is where the judgement lives. It charges each strategy two hundredths of F0.5, so the cascade’s three strategies bring it to 0.78, and the four in the cascade with a learnt model bring that one to 0.77. The cascade wins. One point doesn’t pay for a fourth moving part. Each strategy is another thing to test and maintain, and when a link turns out incorrect, fewer parts mean the fault is found faster.
The charge of 0.02 is mine, and someone else could pick another. Writing it down is the point. Whoever disagrees can see exactly what to argue with, which a decision made in a meeting rarely gives them. The gate and the charge are also the decisions I said someone answers for, so whoever writes them down answers for the links that pass.
Sculley and colleagues describe the machine learning version of this as CACE, “Changing Anything Changes Everything”. They also warn that improving one machine learning model in an ensemble can make the whole system worse when the errors that remain line up with the other components’. They wrote about learnt models, and applying it to a cascade of rules is my own extension, but I think it carries over.
If nothing passes the gate, the task goes back to suggesting links for a person to confirm, or waits. Sometimes the card itself is at fault, and a task that no resolver can do well is really two tasks, or one aimed at a search space that doesn’t hold the answers. None of these answers lasts, since the gold set and the tools both move, so I’d run the trade study again whenever the measurements from part 4 shift.
Where machine learning fits
None of this rules out learnt resolvers. In 2017, Guo, Cheng and Cleland-Huang trained a recurrent network on existing trace links from the Positive Train Control domain and tried 360 configurations of it. It significantly outperformed the classic information retrieval methods. They also retrained it on a larger share of the links, a step a growing project would have to repeat. Lin and colleagues describe deep learning models for tracing as “restricted by availability of labeled data and efficiency at runtime”.
Language models come at the task from the other side. LiSSA, presented at ICSE 2025, recovers trace links with retrieval-augmented generation, prompting a language model with no training on the task itself. It covers three kinds of link, among them architecture documentation to software architecture models, which describe a program’s components much as a systems model describes a system. That removes the retraining cost, although the retrieval index still has to be maintained. It also brings back two questions from this article, whether a hosted language model may read the systems model at all, and whether its fluent explanations are the reasons it really had.
A systems model adds a third. Retrieval finds text that looks like the question, but in a systems model the answer often sits a few links away, in the part that satisfies a requirement or the verification case that verifies it. Can a language model follow the links a systems engineer built?
This is where the hypothetical pipeline leaves the series. It has carried five parts, and part 1 pictured one of its checks failing at two in the morning. It only ever existed in a systems model and in my examples, though, so nothing about it can fail, and no resolver can be tested against it. My demo can. It has a systems model built beside its code, a service that runs, and monitors that report when the service stops answering.
Part 6, coming up next, puts a language model to work on the demo’s own systems model, on my own hardware. It describes how every link the language model proposes, and every reason it gives, will be checked.
Previous: Measuring a resolver · Index: Automating traceability