I used coding agents for months before I admitted how much project knowledge I was leaving in their transcripts.
The files were all there. Claude Code kept JSONL sessions. Codex kept rollouts. Autonomous runs kept plans, iterations, and grader verdicts. In theory I could search them.
In practice, I never did. A transcript full of tool results, file contents, repeated prompts, and command output is a terrible memory system. The useful sentence might be sitting between a 4,000-line build log and a copy of a file that no longer exists.
Scry’s memory domain started with a constraint: keep the source transcripts as evidence, but don’t make the agent read them to remember something.
Distill before asking a model
Raw coding transcripts are mostly tool payloads. Sending all of that to an extraction model would be expensive, slow, and needlessly expose source code and command output.
Scry first runs a deterministic distillation pass in Go.
For Claude sessions, it keeps user and assistant text while replacing tool activity with small breadcrumbs such as [ran tests] or [edited setup.sh]. It records the working directory and repository, then drops the tool request and result bodies.
Codex uses a different JSONL envelope, so its distiller reads session metadata and extracts messages from response events. Autonomous runs are easier because their goals, iterations, gates, and final status are already structured.
What remains is the conversational spine: what we were trying to do, what we decided, and how the run ended. A 4,000-line compiler log stays in the source transcript where it belongs.
Long sessions are split on turn boundaries into chunks of roughly 4,000 tokens with a one-turn overlap. Tiny aborted sessions are skipped. Obvious API keys, bearer tokens, and private key blocks are stripped before the text reaches an extraction provider.
This pass removes most of the volume without asking a model to decide what looks important.
Episodes are the unit of evidence
Each distilled chunk becomes an episode. An episode stores:
- A stable ID derived from the source path and span
- The source type, such as Claude, Codex, autonomous run, seed file, or manual memory
- A reference back to the original transcript range
- When the conversation occurred and when Scry ingested it
- A short summary written by the extractor
The stable ID makes ingestion idempotent. A sweep can encounter the same completed session twice without duplicating its facts.
Scry doesn’t copy the complete transcript into the graph. It keeps the reference. That matters when a summary omits a detail or an extracted fact looks suspicious. The original session remains the record.
Extraction creates entities and facts
The extraction model receives the distilled episode plus a glossary of known entity names and aliases. It returns structured entities and relationships.
Entities have a deliberately small type set: project, service, machine, tool, person, decision, runbook, or concept. They can carry aliases and repository references. “Hermes,” “the Hermes agent,” and an older nickname should resolve to one entity rather than three almost-identical memories.
Facts are edges between entities. A fact might say that a project deploys to a machine, a service uses a model, or one decision replaced another. Each edge includes the plain-language fact, confidence, time bounds, and the episode IDs supporting it.
The extractor doesn’t get to invent a new entity category whenever a noun looks unusual. I learned that from a live failure when an unsupported type stopped the write. It was annoying, but accepting whatever shape the model happened to return would have been worse.
Bad extractions can’t eat the source
The first extraction setup was too optimistic. One model returned an empty reply and left several episodes waiting. Another returned malformed JSON. Model names containing punctuation exposed a separate resolution bug.
None of those failures lost the underlying memories because the source and graph update are different stages.
Scry now supports an ordered extraction chain configured in ~/.scry/config.yaml. If the primary model fails with an empty or invalid response, a fallback can process the same distilled episode. The configuration belongs to the user because cost, privacy, and provider availability differ by machine.
The extraction layer sits behind an interface. Tests feed it sanitized transcript fixtures and canned JSON, then check the resulting graph mutations. Live-provider tests are opt-in.
I don’t expect model extraction to become deterministic. I do expect a bad response to leave the source episode intact and give the fallback model another shot.
Recall is deliberately boring
Scry memory doesn’t use embeddings or a vector database in its first version. Recall searches entity names and aliases, then traverses their facts.
The main operations are straightforward:
scry_recall find entities and their current facts
scry_episodes show sessions connected to an entity
scry_memory_path find the shortest fact chain between two entities
scry_remember store a durable fact deliberately
This works well for the questions I ask most: “what is this project?”, “where does it run?”, “when did we decide that?”, and “how are these two things related?”
Semantic retrieval may become useful if aliases prove too brittle. So far, entity lookup and graph traversal handle the questions I ask. I don’t need to add an embedding stack just because this is called memory.
Memory enters at session start
A memory system that depends on the agent remembering to query it will sit unused.
Claude gets a small orientation block at session start. Codex runs scry memory orient --cwd . before other work. The output stays under a tight token budget: current facts about the repository, a few recently active projects, and a pointer to use full recall for anything unfamiliar.
This is enough to resolve names without dumping the whole graph into context. If the user says “put this on Hermes,” the agent knows Hermes is an existing service and asks Scry before asking the user to define it again.
At the other end, durable decisions are written back with scry_remember. Completed sessions are swept and ingested automatically. The next session gets a compact orientation block and can open the supporting episode if it needs more.
The split between this inferred memory and Scry’s parsed repository data is covered in Why My Knowledge Graph Has Two Halves.