| title | Context as a variable: named, lazy, and queryable |
|---|---|
| description | A precise model for keeping large inputs under stable names, querying them within bounds, and admitting only selected results to model context. |
Audience: people designing or evaluating fak's managed-context runtime who need one precise model for names, paging, filtering, caching, and call results.
Status: this page defines the intended architecture and vocabulary. The lower paging and context-plan seams exist today; the first-class binding/query API is tracked by the linked issues below and is not claimed shipped yet.
“Context as a variable” means a large input lives outside the model's immediate prompt under a stable name. The model operates on that name through bounded queries, and only selected results enter its context window.
outside the model:
tickets -> the complete 500,000-token ticket snapshot
model-visible instruction:
"The source is available as `tickets`."
model action:
count(group(filter(tickets, status == "failed"), owner))
model-visible result:
{"alice": 19, "bob": 7}
The variable is not neural-network hidden state. It is closer to a read-only database relation, lazy collection, array, or file handle in a governed tool runtime.
The precise fak phrase is:
queryable context over addressable sources
“Context as a variable” is the programming-interface metaphor.
Lazy loading is one important part, but not the whole contract. A useful variable has a name and an immutable target; demand may then fetch existing bytes, compute a new view, or admit selected bytes into the prompt. Those are separate transitions:
bind -> fetch/page-in -> materialize/query -> admit
| Stage | What happens | New semantic bytes? |
|---|---|---|
| Bind | A task-scoped name resolves to an immutable source or view identity. | No |
| Fetch / page-in | Existing bytes behind a ref become resident. | No |
| Materialize / query | A filter, projection, or aggregation computes a derived view. | Yes |
| Admit | Selected resident bytes enter the model-visible prompt view. | No new source fact |
| Call | A tool/model recipe executes and creates a result snapshot. | Yes; it may also have effects |
| Refresh | An explicit new execution creates another snapshot and binding revision. | Yes |
Creating or listing a binding must be inert. It performs no source-byte read,
query, call, or prompt admission. A later demand names the work that is needed.
A single loaded boolean cannot represent this lifecycle safely.
Human names are for ergonomics. Machine identities are for correctness.
human alias: tickets
qualified binding: task-42@7:tickets
target kind: call_snapshot
immutable target: sha256:S1
A resolver turns the qualified binding into an exact target kind, immutable identity, policy, and taint record before any fetch, query, cache lookup, call, or admission occurs.
The vocabulary is:
| Term | Meaning |
|---|---|
| Addressable context | Source bytes or a view have a stable immutable/versioned identity. |
| Binding name / alias | Human-facing task-scoped name such as tickets. |
| Qualified binding | Workspace revision plus alias, such as task-42@7:tickets. |
| Context workspace | Task-scoped namespace and versioned binding manifest. |
| Queryable context | Bounded operators compute observations over addressed sources. |
| Derived context view | Immutable, provenance-stamped result of a query. |
| Context program | Policy and plan choosing queries, budgets, residency, and admission. |
| Resident view | Bytes currently available in a serving tier. |
| Admitted view | Bytes actually serialized into the model-visible context. |
An unresolved name such as latest is UI sugar only. It must resolve once to
an exact revision before an operation starts. Cache keys, provenance, and
replay records never use a bare alias as the source identity.
A filter does not mutate its input:
tickets -> source snapshot hash S1
failed = filter(tickets, P1) -> derived view hash V1
failed can itself be named, paged out, restored, filtered again, aggregated,
shared, admitted, or evicted. Its lineage records at least the source snapshot,
canonical query plan, operator version, policy/taint identity, output bounds,
and output hash.
This functional/immutable rule makes filters replayable and makes caching
correct. A cache keyed only by the alias failed would be wrong because an
alias can be rebound in a later workspace revision.
“Cached” is not a sufficient operational explanation. Each layer reuses a different thing:
| Cache | Reuses | Does not prove |
|---|---|---|
| Page/blob cache | Exact source or view bytes already stored. | That a query result is semantically current. |
| Plan cache | A selection/planning decision. | That derived result bytes exist. |
| Derived-view cache | Exact source snapshot + query semantics -> view result. |
That an alias still resolves to the same source. |
| Call/idempotency cache | A witnessed execution outcome for a canonical recipe and scope. | That paging a result should execute a call. |
| Provider KV/prefix cache | Model computation for a stable serialized prefix. | Source truth, tool-result freshness, or query correctness. |
A derived-view key therefore includes every semantic input that could change the answer: all source snapshot hashes, canonical query plan, operator/runtime schema, policy and taint identity, and output-limit contract.
The explain surface must report page_hit, plan_hit, derived_view_hit,
call_outcome_reuse, or provider_kv_hit, not a generic cache_hit.
This architecture reuses fak's existing MMU and plan mechanisms:
internal/ctxmmu/mmu.goreplaces large or held tool-result bodies with governed CAS-backed refs and can restore the same bytes under policy.internal/ctxplan/pagefault.gomodels demand page-fault requests and bounded resolution.internal/ctxplan/materialize.gomaterializes bounded resident views.internal/ctxplan/query.gohas a bounded demand-query selection seam.internal/ctxplan/plancache.gocaches planning decisions, not derived result bytes.
A page-in retrieves bytes that already exist behind an identity. A query creates a new immutable identity. Prompt admission is yet another decision. The named-binding layer joins these seams; it does not introduce a competing paging system.
This is similar to the agent virtual filesystem described in Agent virtual filesystem: both give stable names to content-addressed objects and fault bytes on demand. Queryable context adds typed relational/dataflow operations and model-visible admission semantics. It is also complementary to Addressable KV cache: that page addresses model-computation spans, while this page addresses source and derived data objects.
A tool or model call can produce a large result that becomes addressable context. But a binding must pin a result snapshot, not hide a live call:
call recipe R1 --explicit execution--> call snapshot C1 -> result ref S1
task-42@7:tickets -------------------------------------> S1
Reading, paging, filtering, querying, or admitting tickets uses S1 and
executes zero calls. An explicit refresh adjudicates R1 again and creates a new
snapshot and binding revision:
refresh R1 -> call snapshot C2 -> result ref S2 -> task-42@8:tickets
Only a recipe structurally proven read-only may be deferred until first demand. Effectful calls must never hide behind lazy dereference. A page fault can fetch an existing result blob; it cannot rerun the call that produced it.
Keep three identities separate:
- Call recipe: tool/model, canonical arguments, capability/policy identity, caller scope, and freshness contract.
- Call snapshot: one witnessed execution and immutable result ref/hash, timestamp, taint, and outcome.
- Binding revision: the task-scoped name pinned to that snapshot.
The call/idempotency layer may reuse a witnessed outcome during an explicit refresh. That still differs from the page cache and derived-view cache.
canonical names and immutable resolution (#6533)
|
task-scoped workspace bindings (#6524)
|
demand lifecycle: bind -> fetch -> materialize -> admit (#6531)
| | |
ctxmmu ctxplan query/view (#6518)
|
derived-view cache (#6525)
|
explain and replay (#6528)
|
explicit call recipe -> snapshot -> refresh (#6532)
|
exact aggregation counterfactual (#6526)
|
optional governed helper-model interpretation (#6527)
The implementation order should be:
- Freeze canonical target kinds and resolution rules in #6533.
- Ship the smallest real query/derived-view spine in #6518.
- Add workspaces and inert bindings in #6524, then connect them to fetch/materialize/admit in #6531.
- Add semantic view caching in #6525 and derivation explain/replay in #6528.
- Add immutable call snapshots and explicit refresh in #6532.
- Run the exact aggregation counterfactual in #6526.
- Only then evaluate helper-model interpretation in #6527.
All issues are children of managed-context epic #1570.
It should guarantee:
- binding is inert;
- identities are immutable and replayable;
- page faults never execute calls;
- filters create new views rather than mutate sources;
- each cache has a typed identity and outcome;
- refresh is explicit and versioned;
- source, query, view, call, and admission provenance are inspectable;
- byte, work, storage, call, and prompt-token budgets remain separate.
It does not promise unlimited context or free computation. The runtime must still read or index source bytes; queries consume CPU and storage; admitted observations consume prompt tokens; helper calls consume model tokens. A model can also issue a bad query. The benchmark in #6526 therefore requires exact aggregation answers and the complete cost denominator before any net-true gain is claimed.
- Context management — the current managed-context operator route.
- You never manage the context window — the product promise this architecture should eventually fulfill.
- Context shedding — removing low-value resident turns without invalidating stable cache prefixes.
- Agent virtual filesystem — named, content-addressed objects and demand faults.
- Addressable KV cache — addressability of model computation spans rather than source objects.
- Context as variable vs addressable context study — pinned Prime RLM source study, candidate analysis, and research trail.