Writing /Scry /2026-09-29

Provenance Matters More Than Perfect Recall

Agent memory will get facts wrong. Scry keeps the session, commit, schema, or runtime observation behind every useful answer so an agent can verify it before acting.

Written by
Read time
5 minutes
Topic
Scry

I don’t trust Scry’s memory graph completely.

That isn’t false modesty. A model extracts the memory facts. It can merge two entities that should stay separate, miss a date, or summarize a decision more cleanly than the actual conversation deserves.

I don’t think extraction is going to become perfect. I would rather know where an answer came from and have a way to correct it.

Every memory fact points backward

Scry stores an episode reference with each extracted fact. The episode points at a specific span in a Claude transcript, Codex rollout, autonomous run, seed file, or manual memory.

A recall result includes the current fact and its provenance. If the agent needs more context, it can request the recent episodes touching the entity. The episode summary is still compressed, but its source_ref leads back to the original material.

A low-risk question such as “what is the name of the inference machine?” can use recall directly. A deployment change should inspect the supporting episode and then verify the current machine. The same memory answer can support both workflows because the source is attached.

Without provenance, the graph is just another confident narrator.

Deterministic domains need provenance too

The memory domain makes uncertainty obvious, but parsed graphs can also mislead.

A calls edge may come from SCIP data. A changed_with edge comes from commit co-occurrence. A queries edge may come from a conservative ORM or SQL heuristic. An HTTP edge may be based on one observed request. Those edges don’t deserve identical confidence.

Scry records the source domain and confidence on graph edges. The agent can distinguish a direct foreign key from a relationship inferred because a table name appeared inside a function body.

This matters when a shortest path crosses domains. A path can be structurally valid while one middle edge is weak. Returning only the node names would hide the part most worth checking.

The graph report follows the same rule. High-degree nodes and communities come from stored edges, not an LLM reading the repository and producing a suspiciously tidy architecture essay. Some community labels are awkward. I can still inspect the nodes behind them.

Summaries are indexes, not evidence

Episode summaries exist to help find the right session. They aren’t intended to replace it.

That distinction matters because summaries look more authoritative than they are. A two-sentence account of a long debugging session drops the false starts and uncertainty. It can turn “let’s try this for a week” into something that reads like a permanent architecture decision.

Scry stores the summary beside the source reference rather than treating the summary as the source. The agent can orient quickly, then read the original conversation when the details affect an action.

The same pattern works with git. A commit message indexes intent. scry_intent returns the blame record plus the full commit that introduced the line. It doesn’t assume the short subject contains everything a reviewer needs.

Provenance makes corrections possible

When recall returns a wrong fact, I need more than a delete button.

The episode reference helps identify what went wrong:

  • The source conversation was itself wrong.
  • The extractor misread a correction.
  • Entity resolution merged similar names.
  • A newer session should have superseded an older fact but didn’t.
  • The fact was right at the time and is merely stale.

Those failures need different fixes. I can invalidate a bad extraction, add a missing alias, or mark a relation exclusive so a later fact closes the earlier one. A stale operational fact can be replaced without erasing its history.

If the system kept only the final sentence, every correction would be guesswork.

Confidence is useful when it stays specific

I am skeptical of one global confidence score for an answer assembled from several edges.

Suppose a path says a handler serves an endpoint, queries a table, and was primarily authored by one developer. The route edge may come from observed traffic, the table edge from static analysis, and the author edge from blame over the function’s line range.

Collapsing those into “87% confident” looks scientific and tells the agent very little. Scry keeps confidence at the edge level. The caller can see which relationship is weak.

For memory facts, confidence comes from the extraction response and resolution rules. I don’t use it as a probability that the fact is true. It is a sorting and review hint.

Retrieval quality isn’t the same as memory quality

Agent-memory discussions often focus on whether the correct item appeared in the top few retrieval results. That matters, but it isn’t enough for a system that will change code or operate services.

The retrieved item also needs:

  • A stable identity
  • A time range
  • A source
  • A way to inspect supporting context
  • A correction path

Scry’s entity and edge model is more work than storing chunks and embeddings. It also lets the agent ask how two remembered things relate, what was true at an earlier date, and which session established a fact.

Embeddings may eventually help Scry find entities when aliases fail. They would improve discovery, but the returned fact would still need a source. If recall misses, I search again. If it returns a fact with no source, I don’t use that fact to change anything.

Start a project conversation

Email Jeff

Write your note here, then open it in your email app. This site sends nothing: your email app opens with this draft and you press send there. No email app? Copy the address or the draft instead.

Open in your email app ↗

To jeff@hooton.codes · Closing this window discards the draft.