Skip to main content

Workspace context

The workspace context plane answers questions from what your workspace already contains: the component graph your repos declare, the docs attached to your catalog entities, your repo overviews, and your epic specs. One call returns the passages that answer a question, each citing the entity it belongs to, the repo path it came from, and the commit it was read at.

It is available to agents through two MCP tools, and it holds no truth of its own — every row is derived from the platform's public reads and can be rebuilt from them.

The two tools​

workspace_context​

Ask in plain language when you do not yet know which entity or document you need.

workspace_context(workspace, question, limit?, asOf?)

Returns ranked passages — each with entityRefs, sourceKind, sourceRef, sourceCommit, sourceUri, a score, and a why[] breakdown — plus the neighborhood of every entity cited.

workspace_expand​

Pull the thread once you know the entity.

workspace_expand(workspace, entityRef, depth?, direction?, edgeTypes?, limit?, asOf?)

Walks the component graph outward: out for what an entity depends on, in for what depends on it, both for either. Depth is capped at 3 and an over-deep ask is refused, not clamped — a silently smaller graph looks complete, and that is the one failure you cannot detect. edgeTypes narrows to the relations you care about, and limit sets how many edges come back (default 50, maximum 100); whatever is dropped to keep the response inside its byte budget is reported as candidates-capped rather than quietly omitted.

How the order is decided​

Postgres finds the candidates; the contract ranks them. Search narrowing is a set test over the document's terms, which has no opinion about order. The ordering is a published formula, and this is it:

Your words and the document's are first reduced to stems, so "deploying", "deployed" and "deployment" all count as the same word (Porter's algorithm; names and anything with digits in it are left alone). A passage scores on how many of your words it contains and how close it is in meaning — 60/40 when it has an embedding, your words alone when it does not — and a passage that matches too little of what you asked scores that and nothing else. Once it matches, four published bonuses apply (the term is in the title, the passage belongs to an entity you named, the document was committed recently, and its entity's last run succeeded) which together can add at most 0.5. A passage that scores below the published floor is not an answer, and is not returned.

MatchWeightWhat it measures
Your words0.60Fraction of your query's terms the passage contains
Meaning0.40Closeness of the passage's embedding to your question's

A passage with no embedding is not penalised — it scores on your words at full weight. That is deliberate: it means an index that has not been embedded yet (or an environment where embeddings are switched off) ranks exactly as it did before embeddings existed, rather than uniformly worse.

And the four bonuses on top:

BonusWeightWhen it applies
Title match0.20Your terms appear in the document's title
Named entity0.15The passage belongs to an entity your question named
Recency0.10Decays linearly to zero over 365 days since the commit
Run health0.05The last run of the entity's repo succeeded (unknown counts as nothing)

Run health is repo-level, not per-component: a run is a run of a repo, and a passage inherits its repo's last verdict. That is coarser than "this component's own last build", which is why it carries the smallest weight of the four. An unknown verdict earns nothing rather than half the bonus: a project no run has reported on is the common case, and a term with the same value on every row cannot order anything — it only lifts the floor under passages that should not have cleared it.

The bonuses total 0.5 at most, so they can never promote a passage over one that matches materially more of your terms. Every returned passage carries the why[] that produced its score, so an ordering you disagree with is one you can read rather than guess at.

There is no hidden relevance model, no ts_rank, and no per-tenant tuning. Changing a weight is a versioned contract change, not a configuration knob.

Stemming​

Before anything is matched, both your words and the document's are reduced to stems, using Porter's algorithm. So a question asking how to deploy meets a heading called How deployment works, and "rotating secrets" meets a passage saying there is "nothing to rotate by hand". Without this the term match is exact, and a natural question misses the prose that answers it for no reason a reader could guess.

Stemming reduces word FORMS, not meanings. Asking about "rotating credentials" still will not rank a passage that only ever says secrets — different words for the same idea are a separate problem, and one the semantic half of the score is there to cover.

Two things are left alone, because they are names rather than English:

  • Identifiers — billing-worker, state.run.failed, component:default/api. The identifier is matched whole, and its parts are stemmed, so asking for workers still reaches billing-worker.
  • Anything containing a digit — versions, digests, sha256.

Stemming is a matching key, never something you are shown: why[] reports weights, and passages are returned with their original wording.

Semantic matching​

When embeddings are enabled, a second search runs alongside the keyword one: your question is embedded and the nearest passages are retrieved by vector similarity. That is what lets a question find a document phrased entirely differently — asking about "invoice retries" can surface a runbook that only ever says "dunning".

Both searches only ever find candidates. The order you see is always the formula above, and each passage's why[] shows the two halves separately, so you can tell whether a result came to you by words, by meaning, or by both.

Embeddings are enabled per environment. Where they are off, everything above still works on keywords alone.

What is indexed​

SourceWhere it comes from
catalog_docDocs attached to catalog entities (docs.overview + docs.pages)
repo_overviewA repo's overview.md narrative, via its repo facet
epic_docEpic specs pushed through the tasks plane

Everything is indexed by content digest, so an unchanged document costs nothing to re-index. Nothing is indexed that the public API would not already return to you.

Freshness and honesty​

The index builds itself when you ask. The first question in a workspace is answered straight away from an empty index (flagged index-cold) and starts indexing in the background — ask again a moment later and the answer comes from your documents. After that, a question refreshes the index in the background whenever it is more than a few minutes old, so what you see tracks your pushes without anyone configuring a schedule or a key.

Indexing reads with your own credential and only after checking you can read the catalog, so nothing is indexed that you could not already read yourself. Every response carries indexedAt — how current the picture is — and a degraded[] list when the answer is less than complete:

FlagMeaning
index-coldNothing indexed yet — your question started it; ask again shortly
candidates-cappedThe candidate set was truncated; narrow the question
scope-unreadableThe index could not be read; the answer may be partial
as-of-chunks-approximateasOf was applied to graph edges; passage bodies are present-time

Reads fail soft: a cold or unavailable index returns an empty, honestly labelled pack rather than an error that would break an agent's turn. Authorization does not fail soft — every read is authorized against your own credential, and a workspace you cannot see returns nothing.

Availability​

The context tools appear in the MCP roster only when the plane is deployed and healthy for your environment. When it is not, the roster is exactly what it would have been without it — no tools that exist but cannot answer.