Writing /Scry /2026-09-15

Why My Knowledge Graph Has Two Halves

Scry keeps parsed repository facts separate from inferred memories. The split makes trust, storage, freshness, and failure much easier to reason about.

Written by
Read time
6 minutes
Topic
Scry

The tempting version of Scry was one enormous graph.

Put functions, tables, commits, HTTP requests, projects, machines, people, decisions, and agent sessions into the same store. Give every node a type. Connect everything. Done.

I didn’t build that version. Scry has a deterministic graph for each repository and a separate temporal memory graph shared across projects.

The split adds a little routing work. It also prevents several categories of bullshit, mostly around pretending a model inference is as solid as a foreign key.

A parsed edge and a remembered fact aren’t equivalent

When a SCIP index says function A calls function B, Scry can point at the exact source occurrence. When PostgreSQL says a foreign key connects two tables, the database is the source. When a coverage profile says a line executed, the test runner produced that measurement.

Those are deterministic facts. The indexer may have gaps, but it isn’t interpreting a conversation.

Memory is different. Scry distills a session transcript, sends the compacted conversation to an extraction model, and asks for entities and facts. The extractor may decide that “the mini” and “Hermes Mini” are the same machine. It may infer that a deployment decision superseded an earlier one. Usually that is useful. Occasionally it is wrong.

Putting those edges into the same undifferentiated store would hide the difference. An agent might treat an inferred relationship with the same confidence as a foreign key read directly from the database.

Scry makes the source domain part of the model. Repository graph edges carry confidence and derivation metadata. Memory facts carry confidence, time bounds, and episode references. The boundary stays visible.

Repository knowledge has a different lifecycle

Code changes constantly. A source save can make a symbol index stale within seconds. Scry watches indexed repositories and rebuilds in the background. It writes the new index into a temporary directory, then swaps it into place in about 12 milliseconds.

Git, schema, HTTP, and coverage data each have their own refresh rules. Git history changes when branch refs move. A schema needs another introspection. Captured HTTP requests expire after thirty minutes. Coverage only changes when the user reruns tests and produces a new file.

Global memory moves differently. A decision from May may still be current in September. A project alias should survive a repository rename. A deployment fact should remain queryable after it is superseded because an agent may need to reconstruct an older incident.

Rebuilding memory from scratch every time a source file changes would make no sense. Treating a code index like a historical ledger would be equally strange.

So the stores stay separate. A repository index can be rebuilt and replaced; an old deployment fact needs to remain queryable after a newer fact supersedes it.

The storage boundary follows the security boundary

Repository indexes stay on the machine where the repositories live. That includes code symbols, database metadata, captured local traffic, and graph edges derived from them.

My global memory graph lives on an always-on machine. It is the authority for project names, operational state, and decisions across agent sessions. My laptop reaches it through an SSH-forwarded Unix socket. The memory daemon doesn’t listen on a public or tailnet TCP port.

This arrangement would be awkward with a single physical graph. Sharing memory would also share repository intelligence, whether another machine needed it or not. Keeping the stores separate lets me centralize the knowledge that should follow me while leaving source-derived indexes beside the source.

The MCP layer exposes this split through profiles. A local profile serves code, git, schema, HTTP, graph, and room tools from the current machine. A memory profile serves only recall, remember, episodes, and path queries from the memory authority.

One agent can use both. In normal use, the two sockets are invisible.

Failure should stay inside its domain

Scry’s memory extraction has failed in several entertaining ways. A model returned an empty response. Another response contained invalid JSON. One extraction invented an unsupported entity type. Those episodes stayed available, but the graph update couldn’t complete until the extractor or configuration was fixed.

That shouldn’t affect scry refs.

The same applies in the other direction. A TypeScript indexer missing from PATH should mark the repository partial and leave memory recall alone. The HTTP proxy stopping shouldn’t take down schema queries. A corrupt per-repo graph shouldn’t erase the global record of why the project exists.

Scry uses one daemon and one RPC surface, but the stores and build paths remain independent. I got rid of duplicated daemon plumbing without making a broken transcript extractor take code search down with it.

The join is explicit today

Memory entities have repo_refs, a list of repository paths associated with the entity. When Scry recalls a project, the agent sees those paths and can pivot to the matching repository graph.

That supports a two-step query:

  1. Recall the deployment decision and identify the affected project.
  2. Query the project’s graph for the service, handler, migration, or file involved.

It works, but the traversal happens in the agent. Scry can’t yet run a single shortest-path query from a global decision entity into a function node stored in a repository graph.

I am fine with the manual seam for now. It gives me actual failed queries to design around instead of an excuse to build a distributed graph layer because it sounds neat.

The likely next step is a small cross-graph index, not physically combining the stores. It would map stable memory entities and repository references to code, commit, table, and endpoint nodes. The underlying data could keep its current ownership and lifecycle.

Two halves are still one interface

The user-facing system should feel unified even when the storage isn’t.

An agent starts with a thin orientation from global memory. It uses recall when a project or service name is unfamiliar. Once it enters a repository, it reads the local graph report and uses targeted code, git, schema, or HTTP queries. Durable decisions go back into memory.

That is the current loop. BadgerDB happens to store the data in separate directories, and the MCP profiles route each query to the right socket. Parsed facts and remembered facts meet at the interface, but Scry never labels them as the same kind of evidence.

Start a project conversation

Email Jeff

Write your note here, then open it in your email app. This site sends nothing: your email app opens with this draft and you press send there. No email app? Copy the address or the draft instead.

Open in your email app ↗

To jeff@hooton.codes · Closing this window discards the draft.