Storage and context
Design details for durable transcripts, context preparation, and connected tools.
Status: initial implementation plus follow-up boundaries, 2026-07-31.
This decision originally used the public Claude Managed Agents API as one design reference. Mango now owns its public contract; the storage and execution boundaries below remain because they serve durable, observable, self-hosted agent work. The decision covers Session context, large tool results, Web Search, Web Fetch, MCP, sandboxes, and Temporal.
The central decision is:
Web Search, Web Fetch, and MCP are tools at the public API and policy layers. Their execution owner is an internal capability selected and pinned for each session.
The first implementation may require the configured Messages API base_url to
support native Web Search and Web Fetch. The public Agent schema does not gain a
separate search API key or vendor-specific search configuration. The internal
executor boundary remains in place so a platform-managed search/fetch backend
can be added later without changing the public API or stored session semantics.
Runtime requirements
Mango keeps these runtime requirements:
- A Session is stateful. Its event history and model context survive individual runs, while each Session has an isolated sandbox.
- Agents, Environments, Files, Memory Stores, and Sessions have independent lifecycles. Deleting a Session must not implicitly delete those reusable resources.
- Built-in agent tools include
web_searchandweb_fetch. MCP tools use the same permission-policy concepts as built-ins. - Agent definitions identify MCP servers and enabled toolsets without embedding secrets. Provider and connector credentials belong to deployment/worker configuration, not Agent or Session resources.
- MCP tools default to
always_ask; built-in tools default toalways_allow. Both support opt-inautofor interceptable calls; see Tool permissions. A running Session keeps the tool configuration snapshot with which it began. - Local tool output follows the operator worker's limits. Control-plane MCP output above 100,000 characters becomes a preview and a complete File up to 32 MiB when Files storage is configured; binary MCP content is unsupported.
- Self-hosted sandboxes change the execution location, not the control-plane resource model. Tool inputs and results still cross the control plane.
The raw Messages API adds one requirement that the public event model alone cannot satisfy: native server-tool blocks can contain opaque continuation data, including Web Search encrypted content and citation indexes. Those blocks must be returned to the provider exactly on subsequent calls.
Authority model
There is no single representation that is correct for clients, the model, recovery, and raw bytes. The platform therefore keeps separate logical authorities:
| Question | Authority | Initial physical store |
|---|---|---|
| What did the client observe? | Public Event Ledger | PostgreSQL |
| What exact context continues the model conversation? | Provider Transcript | PostgreSQL JSONB |
| Did a tool possibly change the world? | Operation Journal | PostgreSQL |
| Where are local tool files and processes? | Operator workspace | Operator worker |
| Where are independent public File bytes? | Files API | S3-compatible object storage |
| What knowledge is shared across Sessions? | Memory Store | Separate versioned resource |
| Where are credentials? | Deployment/worker configuration | Environment or operator-managed secret injection |
| What should execute next? | Temporal Workflow state | IDs and small control projections only |
These are logical boundaries, not a requirement for separate services. PostgreSQL initially holds small transactional records and the lossless provider transcript. MCP results are projected into bounded inline content; the control plane does not write raw or binary MCP content into the operator workspace. There is intentionally no general Artifact subsystem in the first implementation. Independent public Files use a narrow S3-compatible byte store without changing the existing tool-result and provider-transcript paths.
Public Event Ledger
The event ledger is the external API truth:
- append-only receipt and commit order;
- the source for event list, replay, and SSE reconciliation;
- stable public event IDs and public tool-use correlations;
- no provider-private blocks, credentials, or unbounded raw output.
It is not the authority for the next model request. Public events are a projection designed for clients and observation. Reconstructing provider context from them loses provider-native blocks, citations, encrypted continuation data, non-text content, and context-compaction decisions.
Provider Transcript
The Provider Transcript is the lossless model-continuation truth:
- stores the exact ordered content blocks accepted from or sent to a provider;
- preserves provider tool-use IDs rather than replacing them with public event IDs;
- preserves unknown provider fields so a client upgrade is not required merely to round-trip a new block;
- records provider, model, server-tool version, and capability-profile version;
- is append-only within an execution attempt.
Blocks are initially stored as JSONB. Hydration must reproduce the provider wire value exactly; a preview or public event payload cannot substitute for it. If provider transcripts later exceed practical PostgreSQL limits, their physical blob storage can change behind the transcript boundary without introducing a public Artifact resource.
A separate mapping relates provider IDs to public IDs:
provider tool-use id <-> internal tool step id <-> public event idThe mapping preserves stable public correlations without mutating the provider transcript.
Context admission and accounting
Every working-model request passes through the same proactive and predictive
admission policy before provider execution. The policy resolves known Anthropic
model windows from the embedded Catwalk catalog, falls back conservatively for
unknown or ambiguous model IDs, reserves the requested output and one
model/tool growth round, and compacts before the provider's input limit. A
provider request_too_large response triggers one more aggressive compaction
attempt for the working turn.
The latest safe provider-usage anchor measures the already-observed request and response prefix. Request and prefix fingerprints invalidate stale anchors, and only messages appended after a valid anchor are estimated. Without a safe anchor, the complete system prompt, messages, and tool schemas use the provider-independent conservative estimator. Outcome and Advisor calls have independent preflight admission, but do not yet share the working turn's provider-overflow recovery path.
Context Snapshot
A Context Snapshot is an immutable recipe for one provider request. It pins:
- the ordered transcript entry IDs included in the request;
- system instructions and effective Agent/Environment revision;
- resolved tool schemas, permission policy, and execution capability profile;
- model parameters and provider adapter version;
- context-policy version and any summary/compaction entry;
- token estimate and parent snapshot.
Compaction creates a new snapshot and summary entry. It never rewrites the original provider transcript. This gives debugging and audit tools an exact answer to both “what happened?” and “what did this model call actually see?”
The current implementation durably checkpoints compacted turn-preparation projections for primary and child Threads: represented transcript event IDs, projected messages, token projection, context-policy version, and the preceding snapshot ID. The Thread's pinned Agent configuration and runtime capabilities remain separately authoritative. Later-round projections within the same turn are not yet stored as equivalent checkpoints. A complete audit recipe for every provider request and attempt, including the resolved endpoint profile, system instructions, tool schemas, adapter version, model parameters, usage anchor, admission limits, compaction reason, and overflow outcome, remains follow-up work; it is not a public Mango resource.
Operation Journal
The existing turn_attempts and tool_steps journal remains the authority for
side-effect recovery:
prepared -> started -> completed
\-> ambiguousAfter started, absence of a durable result is not proof that nothing happened.
An MCP mutation, shell command, custom tool, or platform-managed fetch must not
be retried blindly. Executor-specific idempotency keys can make selected
operations safely retryable, but they do not remove the journal boundary.
Provider calls should gain a similar prepared/started/completed record for
cost, diagnostics, and exact response recovery. A provider idempotency feature
may be used when available; it must not be assumed from an arbitrary
base_url.
Self-hosted workspace, Files, and Skills
The operator-owned sandbox is the Session's mutable execution workspace. It owns processes, intermediate files, tool-created files, and workspace retention. Mango stores durable conversation, Work ownership, approvals, and result correlation; it does not treat the worker filesystem as its state database.
Files uploaded through the public Files API remain independent S3-compatible
objects. A validated UTF-8 File used by user.message or an outcome rubric is
snapshotted before event admission, so later replay does not depend on the
source object. Mango does not automatically mount File or Git Resources into
self-hosted workspaces and does not publish a workspace output directory.
Those transfers belong to the operator launcher.
Custom Skill metadata and immutable Versions keep the split-source design:
PostgreSQL owns identity and pins, while object storage owns the canonical
archive. The Environment worker downloads and verifies the pinned primary and
roster bundles before tool execution. PrepareTurn reads the main SKILL.md
from the same validated canonical archive and projects relative skills/...
paths; supporting files are read from the worker workspace.
Memory Store attachments are the supported Session Resource. The worker synchronizes them through scoped Session APIs and enforces read-only versus read-write roots locally. Large control-plane MCP results are bounded inline rather than written to a server-local path that the worker cannot access.
One tool plane, multiple execution owners
Tool policy and tool execution must be separate concepts.
flowchart LR
Config["Agent tool configuration"] --> Resolve["Session capability snapshot"]
Resolve --> Native["provider_native"]
Resolve --> Managed["platform_managed"]
Resolve --> Worker["client_self_hosted"]
Resolve --> Client["client_custom"]
Native --> Raw["Raw result / provider block"]
Managed --> Raw
Worker --> Raw
Client --> Raw
Native --> Transcript["Exact provider transcript"]
Managed --> Bounded["Bounded MCP content; no file fallback"]
Raw --> Context["Model context projection"]
Raw --> Public["Public event projection"]
Native --> Journal["Operation journal"]
Managed --> Journal
Worker --> Journal
Client --> Journal
The internal execution owners are:
| Owner | Meaning |
|---|---|
provider_native | The configured model endpoint executes a server tool inside the model call |
platform_managed | This control plane invokes a search/fetch service or remote MCP server |
client_self_hosted | The Session parks while the API client executes the built-in in its sandbox/network and returns user.tool_result |
client_custom | The Session pauses and waits for a client-supplied custom-tool result |
The selected owner, provider tool version, capability profile, and permission policy are pinned in the Session runtime snapshot. They are operational facts, not vendor-specific fields on the public Agent resource.
Capability resolution
A URL is an address, not a capability declaration. Even though the first
release requires native Web Search/Fetch support from base_url, the model
adapter must expose an explicit capability profile:
native_web_search
native_web_fetch
native_citations
native_response_inclusion
preserves_unknown_content_blocks
provider_tool_versionsThe profile is derived from a known adapter/configuration, not guessed from the hostname. Agent or Session validation fails closed when an enabled tool cannot be honored. The effective profile is snapshotted so a configuration rollout cannot silently change a running Session.
Web Search
web_search is a built-in tool in the public API. In the first release it maps
to the provider's native server tool.
The adapter must:
- map public
web_searchconfiguration to a pinned, versioned provider server-tool declaration such asweb_search_20260318, rather than exposing it as an ordinary function tool with a permissive input schema; - retain the complete provider response blocks, including citations and opaque encrypted fields;
- pass those blocks back unchanged in later provider requests;
- preserve the distinction between an HTTP/API error and an in-band server-tool error block;
- publish bounded public events linked to, but not substituted for, the exact provider transcript;
- record usage and provider request IDs for diagnostics and billing.
Provider features such as dynamic filtering and response_inclusion are
capabilities, not assumptions. response_inclusion: excluded may reduce
transcript size only where the provider explicitly guarantees that omitted
nested results are not required for continuation. Public operational events and
internal provenance remain durable. Eligibility is decided from the exact
Provider Transcript; an adapter must never exclude a block that a later request
has to resend.
Permission constraint
A native server tool executes inside the provider call. This platform cannot pause between the model requesting the tool and the provider executing it. Therefore:
provider_native + always_allowis supported;provider_native + always_askandprovider_native + autoare rejected during capability resolution;- an interceptable
platform_managedexecutor can supportalways_askby durably parking before execution; client_self_hostedbuilt-ins preserve the permission gate: first wait for a persisteduser.tool_confirmationallowing execution, then execute and returnuser.tool_resultfor the original tool-use event. A denial never authorizes execution, and approval alone does not complete the pending-action barrier.
The provider-native rejection is a current Mango executor limitation. The API
returns a clear unsupported-capability error until Mango has an interceptable
executor. Silently treating always_ask as always_allow is not allowed.
See external tool approvals
for admission and recovery rules.
Web Fetch
web_fetch follows the same native-first execution model and exact-transcript
rules. The adapter retains document blocks, citations, PDF/document content,
provider errors, and cache-control-relevant metadata without flattening them to
text.
A future platform_managed fetch executor must additionally enforce:
- only
httpandhttps, with normalized URLs; - DNS and resolved-IP checks before every connection and redirect;
- denial of loopback, link-local, metadata, private, and disallowed networks;
- redirect count, response size, decompression, media-type, and time limits;
- tenant/domain policy, egress proxy policy, and auditable provenance;
- a deliberate cache and content-retention policy;
- explicit treatment of oversized content and a retrieval contract if full bytes need to remain available outside model context.
The provider's rule that a fetched URL must already appear in conversation context is treated as a security boundary for native execution. A managed executor should enforce an equivalent or stricter policy.
MCP
MCP is a connector subsystem behind the same tool plane, not a special kind of model history.
Configuration and credentials
Reusable Agent configuration stores:
- MCP server name and normalized URL;
- matching
mcp_toolsetenablement and per-tool overrides; - no bearer token, API key, or OAuth refresh token.
This service is agent infrastructure. Agent and Session APIs therefore do not accept connector credentials or references to user-owned keys. Model credentials are read from deployment environment configuration when the worker starts. The current MCP connector supports unauthenticated endpoints only.
If authenticated MCP is added, credentials must remain operator-managed worker configuration, keyed by a deployment-owned connector profile or normalized server identity. Environment variables or the deployment's secret manager may inject the value into the worker, which adds authentication at the outbound transport boundary. Secret values must never enter Agent/Session resources, PostgreSQL event/transcript rows, Temporal payloads, logs, sandbox files, or public errors. Rotation is an operational deployment concern, not a Session mutation.
Discovery snapshot
The MCP connection manager performs initialize and tools/list, validates the
configured toolset, and writes a Session-scoped discovery snapshot containing:
- protocol/server capability metadata;
- tool name, description, and input/output schema digest;
- enabled/disabled decision and permission policy;
- normalized server identity;
- discovery time and snapshot digest.
New tools added by a remote server must not become automatically enabled in an existing Session. The initial implementation should require an explicit allowlist for production Agents even if the public contract permits a broader default.
Streamable HTTP is preferred, with SSE fallback where supported. Connection and authentication failures are nonfatal Session errors: they are observable, leave the Session recoverable, and may be retried on the next transition to running.
MCP invocation
An MCP tool call uses the normal operation journal:
- persist
preparedwith normalized input and discovery snapshot ID; - enforce the recorded permission outcome, parking on
ask, returning a tool error on automaticdeny, or admitting execution onallow; - mark
started; - invoke the remote server with deadlines;
- retain raw MCP JSON in the private journal only when it fits the 100,000-byte diagnostic limit; oversized raw content is omitted;
- create the model and public projections;
- atomically mark the step
completed.
If the connection breaks after started, the step is ambiguous unless the
tool has a proven idempotency contract. The runtime must not infer that read-like
tool names are safe.
MCP text, textual embedded resources, links, and structuredContent become
model-visible text. Image, audio, and binary resource content become an explicit
unsupported-content message. _meta and protocol control fields stay outside
model context; bounded raw JSON may remain in the private journal.
If the projected text exceeds 100,000 characters and fits within 32 MiB,
Mango retains the complete text as a regular File in configured object storage.
The model and public event receive a 2,000-character preview with a path relative
to the Session workspace root: .mango-tool-results/{file_id}.txt. The public
agent.mcp_tool_result also carries file_id. Native Go workers materialize
these Files before owned local tool dispatch, including after paginated history
replay and Work reclaim. Use Bash or a bounded file read to inspect selected
parts; the ordinary read tool still has its 64 KiB limit.
The tool journal commits its complete private byte receipt before object-store I/O.
JSON base64 encoding preserves every UTF-8 byte, including NUL, in PostgreSQL
JSONB. Previews and inline text represent NUL as the printable \u0000 escape;
native raw JSON containing that escape is omitted from diagnostics.
An Activity retry publishes the same File without repeating the MCP invocation.
After File publication, the private full text is removed from the journal; only
the preview and reference reach Temporal history. File downloads verify size and
SHA-256 and publish through a confined worker filesystem root. _meta and binary
content remain outside the saved textual projection.
Generated Files are pinned while the owning Session exists. Active scoped Work credentials can read only their own Session's output Files, without list, upload, or delete access. Workspace keys can download them normally. Session deletion releases the File pin without deleting bytes; operators then use the ordinary Files deletion API. Startup reconciliation leaves resumable generated uploads alone while their Session exists and cleans abandoned uploads after its deletion. A publication retry cannot recreate a File after Session deletion.
Without configured Files storage, or above the 32 MiB projected-text limit,
Mango returns a tool error explaining that full retention was unavailable and a
bounded preview. It does not report successful full retention or a usable path.
A configured object-store outage retries publication before advancing the turn.
External workers should implement the documented file_id download contract;
Mango does not reach back into their filesystems.
The first MCP slice supports tools only. Resources and prompts should be added only when Mango's product requirements and context policy define their lifecycle.
Provider-round transaction model
Intermediate model/tool rounds must survive Activity retries while preserving their public receipt order.
- A turn starts from the committed Provider Transcript.
- The Workflow durably appends
span.model_request_startbeforeCallModel. CallModelretains the complete provider response blocks in the turn's private transcript delta.- A tool Activity journals execution and adds its model projection to that delta.
- Before another provider round begins, its predecessor's completed public model/tool events are appended idempotently with deterministic IDs.
- Turn completion atomically commits the remaining public events and transcript delta, marks the trigger processed, and updates Session status.
- A later turn loads the private transcript instead of reconstructing provider context from public events.
This model exposes completed progress without allowing a later model request to overtake it, while keeping provider-native conversation state lossless.
Temporal boundary
Temporal manages ordering, timers, retries, cancellation, approval waits, Continue-As-New, and cleanup sagas. It is not the transcript or file store.
The target Workflow and Activity payloads carry:
- Session, run, attempt, provider-round, tool-step, and Context Snapshot IDs;
- small status/control projections;
- digests and bounded error summaries.
They should not carry full model requests, provider responses, MCP results, fetched documents, or file bytes. The current implementation still records the bounded request and transcript delta in Temporal history while PostgreSQL atomically commits it. Moving to context/round IDs is follow-up work. File bytes already stay in the sandbox and never enter Temporal payloads.
Records
Names are illustrative; they may share the existing PostgreSQL service:
| Record | Key fields |
|---|---|
provider_transcript_turns | Session, canonical trigger, all represented resolution event IDs, ordered lossless message delta, public/provider tool ID mappings |
tool_steps | prepared/started/completed/ambiguous state, bounded raw result or sandbox path, model projection |
mcp_discovery_snapshots | Session/server, normalized URL, immutable discovered tool definitions |
Every table is tenant-scoped. Sensitive content is encrypted at rest, access is audited, and public reads authorize through the owning resource rather than accepting raw storage keys.
Lifecycle and deletion
| Resource | Session deletion behavior |
|---|---|
| Public events and provider transcript | Delete according to Session retention contract |
| Session sandbox | Durably tear down through the existing deletion saga |
| Session-owned tool-result files | Delete with the Session sandbox |
| Independent uploaded File resources | Preserve |
| Memory Stores and immutable versions | Preserve |
| Agent and Environment definitions | Preserve |
| Deployment credentials | Outside Session lifecycle; rotate through deployment operations |
| Attempt/journal records | Retain for the bounded audit window, then purge with the Session |
Deletion first fences new work, terminates orchestration, tears down the sandbox, and then removes or tombstones authoritative records. Independent File deletion follows the File resource's own lifecycle.
Implementation status
The current code has a public ledger, outbox, Temporal workflow, Environment Work leases, and tool ambiguity journal. This change adds the minimum context and tool boundaries needed for native web and unauthenticated MCP:
- Committed turns load a lossless Provider Transcript rather than reconstruct provider context from public events.
- The provider wire model retains opaque response blocks and replays them unchanged.
- Provider tool-use IDs remain private; explicit mappings connect them to public event IDs.
- Local tools enforce worker-side output limits; oversized MCP projections use a bounded preview and do not create a sandbox file.
- Native Web Search/Fetch declarations go to the configured Messages API
base_url; provider-native web rejectsalways_askandauto. - Remote MCP tools use the official Go SDK, Session-pinned discovery, normal
permission policies, and the operation journal. Discovery connection
failures emit a recoverable
session.error, omit that server's tools for the turn, and are retried on a later turn. The default connector resolves and pins public IPs per connection and rejects loopback, link-local, private, metadata, and reserved networks; private MCP requires a future explicit tunnel/egress capability. - Request-time token-aware context projection and extractive compaction are deeply detached from the durable transcript, including nested tool inputs and rich/raw content, so request adaptation cannot mutate stored history. Known model windows come from Catwalk; provider usage anchors measure the observed prefix, predictive admission reserves the next round, oversized tool results are compacted independently, and a working turn can recover once from provider-reported context overflow. Compacted turn-preparation projections are durably snapshotted per Thread. Complete per-request audit recipes, later-round projection checkpoints, explicit custom-endpoint capability overrides, equivalent Outcome/Advisor overflow recovery, deployment-managed MCP authentication, provider-request records, and reference-only Temporal payloads remain follow-up work.
Delivery order
- Context foundation: Provider Transcript, lossless provider blocks, and provider/public ID mappings. Implemented for committed turns.
- Payload foundation: Local worker output limits and bounded inline MCP projections. Full MCP result transfer is not implemented.
- Native web: provider capability profile plus native Web Search/Fetch in
always_allowmode, with exact replay, citations, and a hard per-request context-size ceiling. Native declarations and replay are implemented; explicit per-endpoint capability configuration remains. - MCP tools: Agent/server validation, discovery snapshots, approval parking, and journaled invocation. Unauthenticated Streamable HTTP, discovery snapshots, approval, and invocation are implemented; deployment-managed authentication remains.
- Self-hosted execution: built-in calls park for
user.tool_resultand are implemented. Optional managed search/fetch providers remain follow-up work. - Context engineering: Catwalk-derived model windows, provider-usage anchors, conservative post-anchor estimates, predictive admission, rich-content-aware projection, extractive and oversized-tool-result compaction, one-shot working-turn overflow recovery, and immutable per-Thread turn-preparation checkpoints are implemented. Explicit custom-endpoint overrides, complete per-provider-request audit recipes, later-round projection checkpoints, equivalent Outcome/Advisor overflow recovery, provider-exact counters, compaction quality evidence, and retention controls remain follow-up work. Independent cross-Session Memory now uses PostgreSQL-backed Stores and Docker Session mounts.
The first two steps were treated as prerequisites rather than cleanup after native web, so new Sessions do not depend on reconstructing provider context from flattened public events.
Historical and implementation references
These references informed parts of the original design. They are provenance, not Mango's target contract or roadmap. Where Mango retained a route, resource, event, or field shape, that surface is now governed by Mango's documented semantics and may evolve independently.