[docs] Plan agent multi-modality#5439
Conversation
Design workspace for multi-modal agent input (images, audio, documents): the model perceives the content and the agent can operate on it, with every shared file a durable, findable record. Docs only. README, context (verified current-state audit), proposal (design + 8 interaction diagrams + decision log), status.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Important Review skippedNo new commits to review since the last review. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
📝 WalkthroughWalkthroughAdds a documentation set for agent multi-modality attachments, covering current limitations, target storage and delivery architecture, staged implementation, capability constraints, decisions, status, scope, and validation requirements. ChangesAgent multi-modality design
Estimated code review effort: 2 (Simple) | ~10 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 6
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 6e6cffaa-1cf6-4746-a5d0-51e72e498e22
📒 Files selected for processing (4)
docs/design/agent-workflows/projects/agent-multi-modality/README.mddocs/design/agent-workflows/projects/agent-multi-modality/context.mddocs/design/agent-workflows/projects/agent-multi-modality/proposal.mddocs/design/agent-workflows/projects/agent-multi-modality/status.md
| - **The original** — *immutable, agent-unreachable.* Uploaded once to a session-scoped | ||
| **attachments mount** that is **never geesefs-mounted into the sandbox**. It is the source of | ||
| truth: what the model reads, what the drawer lists, what download returns. Nobody edits it. | ||
| - **The working copy** — *mutable, in `cwd`.* The runner materializes referenced files into the | ||
| agent's working directory. The agent may read, transform, edit in place, or delete it. Its | ||
| edits become new artifacts; the original is untouched. |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Enforce the “immutable original” invariant.
Keeping the attachments mount out of the sandbox prevents agent access, but the proposal still reuses a write-by-destination-path upload flow. Define server-side write-once semantics—or immutable object IDs/versioning—and reject overwrite/delete operations. Otherwise the client or another API path can mutate the source of truth, breaking download and findability guarantees.
| Key change is step 12–13: the runner resolves the reference and emits **real ACP blocks** | ||
| instead of `[{type:"text", text: turnText}]`. Steps 8–10 (materialize working copy) run once per | ||
| session per file — `cwd` is durable across turns, so later turns skip it. |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
Define re-materialization after working-copy deletion.
The design allows the agent to edit or delete the working copy, but later turns “skip” materialization unconditionally. Reusing the reference after deletion can therefore leave the agent without the file. Specify an existence check and rehydration policy, while ensuring existing edits are not overwritten unexpectedly.
| The runner already receives ACP `promptCapabilities` at initialize and ignores them | ||
| ([context.md](context.md)). We read them, map them to Agenta capability flags, return them to the | ||
| FE, and gate both ends. | ||
|
|
||
| ```mermaid | ||
| flowchart TD | ||
| A["Harness init (ACP)"] --> B["AgentCapabilities.promptCapabilities<br/>{ image, audio, embeddedContext }"] | ||
| B --> C["runner maps → Agenta capabilities<br/>{ images, audio, documents }"] | ||
| C --> D["surfaced to FE with model/harness meta"] |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Specify the capability contract across every layer.
The supplied runner code currently defines images and fileAttachments; this proposal introduces images, audio, and documents. Explicitly update the runner flags/types, Agenta protocol/DTOs, API response, and FE metadata, including the embeddedContext → documents mapping. Otherwise the composer and runner cannot enforce the same gate.
| subgraph run["runner enforcement (defense in depth)"] | ||
| F --> J{capability present at prompt build?} | ||
| G --> J | ||
| H --> J | ||
| J -->|yes| K["emit ACP block"] | ||
| J -->|no| L["drop with explicit warning event<br/>(never silent)"] | ||
| end |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Reject unsupported modalities instead of dropping them.
“Drop with an explicit warning” still lets the run continue without content the user attached, violating the stated “never silently dropped” and honest-capability goals. Fail the turn with a user-visible structured capability error, rather than invoking the model with a partial prompt.
| | --- | --- | --- | | ||
| | Composer picker / drag / paste / preview | [`web/oss/src/components/AgentChatSlice/assets/files.ts:21`](../../../../../web/oss/src/components/AgentChatSlice/assets/files.ts) | ✅ base64 `data:` URL as AI-SDK `FileUIPart` | | ||
| | Request body `data.inputs.messages[].parts[]` | [`web/packages/agenta-playground/src/state/execution/agentRequest.ts:406`](../../../../../web/packages/agenta-playground/src/state/execution/agentRequest.ts) | ✅ | | ||
| | Vercel adapter parses the part | [`sdks/python/agenta/sdk/agents/adapters/vercel/messages.py:97`](../../../../../sdks/python/agenta/sdk/agents/adapters/vercel/messages.py) | ✅ → `ContentBlock(type="image"|"resource")` | |
There was a problem hiding this comment.
what are the contentblock types in the vercel adapter?
There was a problem hiding this comment.
are there limits we should know about
There was a problem hiding this comment.
🤖 The AI agent says:
The Vercel adapter (sdks/python/agenta/sdk/agents/adapters/vercel/messages.py) produces five content block types: text, image, resource, tool_call, and tool_result. For an incoming file part the rule is: a media type starting with image/ becomes an image block, and everything else (PDF, audio, any document) collapses into a resource block. There is no audio block today.
Changed in the PR: research.md section 10 documents the adapter's full vocabulary, the file rule, and the missing audio type. Adding a first-class audio block is part of the plan.
There was a problem hiding this comment.
🤖 The AI agent says:
Yes, two kinds of limits. Our own front-end limits (5 files, 5 MB each, image/PDF/text/JSON) are placeholders; the code comment on them says they were meant to be replaced with limits derived from the selected model. The real limits come from the providers: Claude caps a request at about 32 MB, image size limits depend on the deployment, and PDF support goes up to 600 pages (100 pages on 200k-context models). Gemini has its own per-request limits.
Changed in the PR: research.md section 3 lists the provider limits with sources, section 10 covers our placeholders and where they came from, and the plan derives the front-end limits from the selected model instead of the placeholders.
|
|
||
| Three concrete failure modes: | ||
|
|
||
| 1. **Image + text** → the model gets only the text. On cold replay it gets the literal string |
There was a problem hiding this comment.
pleaes let us not limit ourselves to image in this research. let us be as general as we can. images, zips, cad files, excel files, folders, videos.. different sizes etc..
There was a problem hiding this comment.
🤖 The AI agent says:
Done. research.md section 3 now covers images, audio, video, PDFs, office documents (Excel, Word), plain text and code, CSV and JSON, zip archives, CAD and other binary formats, folders, and very large files.
The organizing idea that came out of it: a model natively perceives only a small set of kinds (images, audio for some models, documents for some, video only for Gemini). Everything else is a file the agent opens with tools, sometimes after converting it (a spreadsheet through a library, a zip by extracting). The design therefore puts every attached file on the agent's disk for tool use, and additionally delivers the native kinds to the model inline when the model supports them.
| Capability flags were designed and abandoned mid-flight: | ||
|
|
||
| - Agenta flags exist in TS ([`protocol.ts:283`](../../../../../services/runner/src/protocol.ts) — | ||
| `images?`, `fileAttachments?`) and Python ([`dtos.py:187`](../../../../../sdks/python/agenta/sdk/agents/dtos.py)). |
There was a problem hiding this comment.
why do we have two types only / imges and fileattachements
There was a problem hiding this comment.
🤖 The AI agent says:
Because the flags were built ahead of the feature that would use them. The whole capability set, including images and file_attachments, landed with the original agent-runtime protocol commit (b9e62f99aa). The docstring even says adapters should "skip image blocks without images", but that skip was never built, because nothing sends image blocks yet. Only the tool flags (mcp_tools, tool_calls) ever got a real gate.
Changed in the PR: research.md section 9 has this history, and the design replaces the two-flag set with a layered capability model (protocol, adapter, model) described in design.md.
| - Agenta flags exist in TS ([`protocol.ts:283`](../../../../../services/runner/src/protocol.ts) — | ||
| `images?`, `fileAttachments?`) and Python ([`dtos.py:187`](../../../../../sdks/python/agenta/sdk/agents/dtos.py)). | ||
| - `dtos.py:187` documents the intent — adapters *"skip image blocks without `images`."* | ||
| **No such skip exists.** `capabilities.images` is assigned once at |
There was a problem hiding this comment.
again why capabilities are abut images only
There was a problem hiding this comment.
🤖 The AI agent says:
Same root cause as the previous comment: the runner probes the harness, maps only images and fileAttachments out of the ACP capability set (image, audio, embeddedContext), and then never reads them anywhere.
Changed in the PR: the design now requires the runner to map all three ACP capabilities, and treats the effective capability for a modality as the intersection of three layers: what the protocol can carry, what the harness adapter actually delivers natively, and what the selected model perceives. The layered contract table is in design.md.
| | Type | Owner | Modalities | Where | | ||
| | --- | --- | --- | --- | | ||
| | ACP `ContentBlock` | Zed (external) | text, image, **audio**, resource_link, resource | runner → harness ([`run-turn.ts:478`](../../../../../services/runner/src/engines/sandbox_agent/run-turn.ts)) | | ||
| | Agenta `ContentBlock` | ours | text, image, resource — **no audio** | FE → runner ([`protocol.ts:12`](../../../../../services/runner/src/protocol.ts), [`dtos.py:229`](../../../../../sdks/python/agenta/sdk/agents/dtos.py)) | |
There was a problem hiding this comment.
what is a ressource / ressource_link
why ours have no audio.
please update this whole thing and explain the different modalities and how they are usually handled. which ones are special which one are just files
There was a problem hiding this comment.
special as in native and require special path
There was a problem hiding this comment.
🤖 The AI agent says:
Definitions first. An embedded resource carries the file's bytes inline in the message (as text or base64), so the harness can hand the content to the model. A resource_link is only a pointer (a URI and a name); the agent may fetch it with a tool, but the model is never guaranteed to see it. Our content block simply has no audio variant today (audio collapses into resource); the plan adds one, because ACP represents audio as its own content type and there is no tool-read fallback for it.
Changed in the PR: research.md section 3 is now a full taxonomy of modalities: the native kinds, where the model's own encoder must receive the bytes through message content (images, audio, documents for some models, video for Gemini), versus everything else, which is just a file the agent opens with tools. Section 4 explains each ACP content type and what our two pinned harness adapters actually do with each one.
There was a problem hiding this comment.
🤖 The AI agent says:
Exactly. "Special" means native: the model's own encoder must receive the bytes through the message content, and no tool that reads a file from disk can substitute for that path. research.md section 3 uses this as the dividing line: native kinds need the dedicated delivery path, everything else is a normal file for the agent's tools.
| external spec and the harness already support image and audio; the whole build is on **our** side | ||
| of line 478. | ||
|
|
||
| ## The substrate we build on: mounts |
There was a problem hiding this comment.
please add a research file research.md using the plan-feature skill and make sure to explain exactly the capabilities / endpoints we have in mounts. who has the right to use them. what information are required to use them (arguments). this will help us understand the flows in which they would be used. for instance how they would be used in the drawer (to updload / edit / remove / move files within the mount) or in the message and who has the responsibility to do things. same for the previous context we need research about the different modalities and which ones are special. research about how acp works with these modalities. research about how zed deal with these in acp, about opencode how it deals with these (especially in the ui) and any opensource solution that might have this. to understand best practices flows and responsabilities. also we need to udnerstand capabilities of models and what makes things special
There was a problem hiding this comment.
🤖 The AI agent says:
research.md now exists and covers all of this.
Mounts: section 2 has a table of every endpoint (route, method, required permission, arguments, what it does), covering create, query, sign, upload, write, download, list, delete, folder creation, and export. Two facts worth calling out: there is no move or rename endpoint (a move is a write to the new path plus a delete of the old one), and there are two separate byte paths (through the API with permission checks, or directly against S3 with signed credentials, which today are always read-write). Section 2 also maps the endpoints to the three flows and who owns each: the drawer (front end calls upload, write, delete, list; the API enforces permissions), the message attachment (front end uploads once, API stores and validates, runner resolves), and materialization (runner writes the working copy).
Modalities: section 3 (which are native, which are just files, provider limits). ACP: section 4 (each content type, what is required, and what our two pinned adapters actually do with each block). Zed and opencode: section 6 (how each collects files, how they reach the model, who enforces limits). Model capabilities: section 3.
| ([`service.py:146`](../../../../../api/oss/src/core/mounts/service.py) `_is_internal_mount_path`, | ||
| `service.py:182` `_is_hidden_path`). | ||
|
|
||
| ## Side effects the current design already causes (fixed for free by the reference model) |
There was a problem hiding this comment.
you are conflating context, research, proposal and design. please reorganize the research in a way that is readable by a human being
There was a problem hiding this comment.
🤖 The AI agent says:
Restructured. The folder now follows the plan-feature layout: context.md is the plain story only (what happens today, what we want), research.md holds all findings organized by question, design.md is the design written as options and decisions with the diagrams explained in prose, plan.md is the staged work, scope.md is in and out of scope, decisions.md is the compact decision log plus open questions, and status.md tracks state. proposal.md is deleted; its content was split into design.md and plan.md. README.md opens with a glossary and a reading order.
|
|
||
| ## Side effects the current design already causes (fixed for free by the reference model) | ||
|
|
||
| - **History resend amplification.** The full message history is resent every turn |
There was a problem hiding this comment.
this is a limitation pre-session work. after sessions we will send only the last message. even in case of catastrophic failure the responsability of sending all messages will go to the runner
There was a problem hiding this comment.
basically in case the session cannot be continued the runner should read from records and then rerun them (let's create an issue for this in case we did not already)
There was a problem hiding this comment.
🤖 The AI agent says:
Agreed, and it is now framed that way. research.md section 7 describes resending the full history as the pre-sessions behavior, and states the direction: the front end sends only the last message, and the runner owns rebuilding context from records when a session cannot be resumed warm.
One verified fact that matters for the plan: records today are text-only (the record schema allows only {type: "message", text}), so carrying attachment references in records requires extending the record schema. The plan makes that extension part of the first release, so the records are ready for the runner-rebuilds direction.
There was a problem hiding this comment.
🤖 The AI agent says:
There was no existing issue, so I created one: #5443, "Runner should rebuild session context from records when a session cannot be resumed warm". It links #5384 (the warm case: hold client tool calls open instead of replaying) and this PR (records will carry attachment references, which a faithful rebuild needs). research.md section 7, scope.md, and status.md link to it.
| - **History resend amplification.** The full message history is resent every turn | ||
| ([`agentRequest.ts:401`](../../../../../web/packages/agenta-playground/src/state/execution/agentRequest.ts)); | ||
| a 5 MB image is ~6.7 MB base64 re-uploaded on *every* subsequent turn — quadratic. | ||
| - **localStorage eviction bomb.** Messages persist with base64 inline |
There was a problem hiding this comment.
please always write using context (update for the whole document)
There was a problem hiding this comment.
🤖 The AI agent says:
Rewritten in plain words. What it means: saved chat messages currently keep the file's bytes inline, encoded as text. The browser gives a site about 5 MB of local storage. One attachment can exceed that, and when the save routine runs out of room, it deletes other sessions' saved chats to make space. So attaching one file in one conversation can silently erase the saved history of other conversations.
Changed in the PR: context.md and research.md section 1 explain it this way, and the fix falls out of the design: once messages store a small reference instead of the bytes, the storage pressure is gone.
There was a problem hiding this comment.
🤖 The AI agent says:
Applied across the folder. README.md now carries a glossary of the domain terms the files lean on (runner, harness, adapter, ACP, mount, composer, materialize, records, attachment resource, warm and cold sessions, and so on), the main flows introduce each entity together with which system owns it and where it runs, and the known coinages were removed or replaced with plain words. If you still hit a term that is used before it is explained, flag it and I will fix that spot.
| ([`sessions.ts:106`](../../../../../web/oss/src/components/AgentChatSlice/state/sessions.ts)); | ||
| one attachment blows the ~5 MB budget and `writeMessagesWithQuotaGuard` responds by deleting | ||
| *other sessions'* transcripts (`sessions.ts:409`). | ||
| - **Tracing payload bomb (Python side).** `_agent` is instrumented with |
There was a problem hiding this comment.
this should be taken into accoun tin the plan
There was a problem hiding this comment.
🤖 The AI agent says:
Done. plan.md folds the tracing change into the first release: once messages and records carry a reference instead of bytes, the Python agent span (which keeps messages on purpose and has no length cap) holds a small reference rather than a base64 blob. The release verifies that outcome instead of adding a truncation cap, since the cause is removed at the source.
|
|
||
| Once history and traces carry a **reference** instead of base64, all three dissolve. | ||
|
|
||
| ## Scope confirmed with the product owner |
There was a problem hiding this comment.
I would add also large copy pasta. At least as a follow up. If the user copy pastes a large document, we should not send it though the prompt, it should instead be sent as a file (question where would it be saved maybe answered elsewhere) and then referenced in the prompt instead.
There was a problem hiding this comment.
when we upload something we should show the reference in the prompt in a nice way (out of scope probably but follow up)
There was a problem hiding this comment.
🤖 The AI agent says:
Added as a follow-up in scope.md: when a person pastes a very large block of text, it should be saved as a file and referenced in the prompt, the same path an attachment takes. On your "where would it be saved" question: the attachments mount this project builds is the leading option, since the reference mechanism would be identical, but the follow-up still needs its own storage decision when it is picked up.
There was a problem hiding this comment.
🤖 The AI agent says:
Added as a follow-up in scope.md: a clean reference chip in the composer (filename and kind) when a file is attached, as a presentation improvement on top of the core flow.
|
|
||
| ## Goals | ||
|
|
||
| 1. **Model perceives it.** Image, audio, and document content reaches the model as standard ACP |
There was a problem hiding this comment.
why are we saying that they should reach the model as standard acp blocks. what are the advantages here instead of dealing with it ourselves (that is uploading to the mount then referencing).
Please add research to explain how does ACP uses these blocks. what happens under the hood when communicating with the harness. which harnesses actually work/use these acp blocks. and how difefrent is it from the other approcah i mentioned
There was a problem hiding this comment.
🤖 The AI agent says:
The research now answers this with the actual adapter code, and the answer is: we need both paths, because they do different jobs.
Under the hood: an ACP image or audio block carries required base64; the harness adapter converts it into the model API's native content. We verified the two adapters we actually pin: the Claude adapter (@agentclientprotocol/claude-agent-acp 0.58.1) converts ACP images into native Claude image blocks, and Pi (pi-acp 0.0.29) passes images through as image inputs. So inline ACP blocks are the reliable way to make the model see an image.
The alternative you describe (upload to the mount, mention the path) gives the agent tool access but not reliable perception: ACP has no "model reads a URL" content type, and Claude Code's Read tool does not dependably return an image as vision input. For audio there is no tool fallback at all.
So the design does both: upload to the mount and reference (for storage, findability, and the agent's working copy) plus inline delivery at prompt time (for perception, rebuilt by the runner from the stored original). research.md sections 4 and 5; design.md decision D1.
| 2. **Agent operates on it.** The file is on the agent's filesystem so its tools can read, | ||
| transform, or edit it — whatever the conversation calls for. | ||
| 3. **Always findable.** Every file a user shares is a durable, immutable record surfaced inline | ||
| in the conversation and under a "Shared by you" origin in the Files drawer — never lost, even |
There was a problem hiding this comment.
not sure about never lost. if the file is uploaded this means it can be deleted
There was a problem hiding this comment.
🤖 The AI agent says:
You are right; the wording overclaimed. The guarantee is now stated precisely: the agent cannot change or delete the original, because the attachments mount is never exposed in the sandbox and originals only enter through a create-only upload route (the API refuses overwrite and delete for originals, and no signed storage credentials are issued for that mount). It is not immortality: Stage 1 defines cleanup for uploads that were never sent, and what happens on session archive or delete is an open retention question in decisions.md. design.md decision D7.
| 3. **Always findable.** Every file a user shares is a durable, immutable record surfaced inline | ||
| in the conversation and under a "Shared by you" origin in the Files drawer — never lost, even | ||
| after the agent runs. | ||
| 4. **Honest capability.** Unsupported modalities are refused at the composer, driven by the ACP |
There was a problem hiding this comment.
what does it mean to have unsopported modalities? an agent using an llm that does not support images can still use tools that use images. no? please explain in research and address this comment
There was a problem hiding this comment.
🤖 The AI agent says:
Correct, and this comment changed the design. Perception and tool use are now treated as two separate things. "Unsupported modality" only ever means the model cannot natively perceive that kind. The file still lands on the agent's disk, so tools work regardless of the model.
The gating now works accordingly: attaching a file always puts it in the agent's workspace; native delivery to the model happens when the capability intersection (protocol, adapter, model) allows it; when it does not, the composer says so up front and the turn proceeds with the file as a workspace-only attachment. The runner fails the turn only when it is explicitly asked to emit a native block the harness cannot accept. research.md section 3; design.md decisions D5, D6, and the open decision D11 (what attaching promises in the first release).
| after the agent runs. | ||
| 4. **Honest capability.** Unsupported modalities are refused at the composer, driven by the ACP | ||
| `promptCapabilities` the runner receives — never silently dropped. | ||
| 5. **No base64 on the hot path.** The wire, the persisted history, and traces carry a small |
There was a problem hiding this comment.
please be very precise here. what do you mean by hot path, what do you mean by history. to my knowledge we dont have an entity called history. and why do we care about the wire between the service and the runner not having base64, what is the value there?
There was a problem hiding this comment.
🤖 The AI agent says:
Rewritten without those words. Precisely: the wire is the request from the front end through the API to the runner. There is no entity called "history"; the two real things are the front end's saved message store (browser local storage) and the durable per-session records kept by the API. The docs now name those.
Why carrying base64 there hurts: the front end currently resends the whole conversation on every turn, so the same image bytes travel again on every later turn (the cost grows with the square of the conversation length); the saved bytes overflow browser storage and evict other sessions' chats; and the bytes land verbatim in traces. A small reference removes all three. That is the value, and it is separate from the model boundary, which still needs inline bytes at prompt time. research.md section 1.
|
|
||
| ## Non-goals (for now) | ||
|
|
||
| - **Video.** Absent from ACP's content model and from every model's metadata; net-new, deferred. |
There was a problem hiding this comment.
again not sure i understand the acp content model and what does it actually mean for us
There was a problem hiding this comment.
🤖 The AI agent says:
research.md section 4 now lays it out concretely: the five ACP content types, their exact fields, and which fields are required. What it means for us is one fact: for the model to perceive a file, the bytes must be inline in the turn at the moment the runner calls the harness. Storage never removes that requirement; it only removes bytes from the wire and the saved state. And one caveat we verified in the pinned adapters: what the model actually receives depends on the adapter, not just the protocol (images arrive natively; documents and audio do not today).
| ## Non-goals (for now) | ||
|
|
||
| - **Video.** Absent from ACP's content model and from every model's metadata; net-new, deferred. | ||
| - **Assistant-emitted files.** A runner `file` run-event type is declared but has no emitter |
There was a problem hiding this comment.
can you explain in the research what are these assistant emitted files. how the flow look like there. where is this even used. i would expect an agent to just write the files to the mount and reference them in its answer (by the way how does the frontend know when to show a link to a file [in agenta but also in zed])
There was a problem hiding this comment.
🤖 The AI agent says:
research.md section 8 covers this. The runner protocol declares a file event ({type: "file", url, mediaType}), but no runner code emits it. Your expectation matches how Zed and opencode actually behave: the agent writes files into its working directory and mentions them in its answer, and the client shows the file because it can see the working directory, not because of a protocol event. In Agenta the drawer already lists cwd, so an agent-produced file is already findable there; the missing piece is an inline "here is the file" chip in the answer, which is separate work and listed in scope.md as out of scope here.
| - **Video.** Absent from ACP's content model and from every model's metadata; net-new, deferred. | ||
| - **Assistant-emitted files.** A runner `file` run-event type is declared but has no emitter | ||
| ([context.md](context.md)); making the agent *return* generated files is separate work. | ||
| - **Content-addressed dedup / thumbnailing / transcoding.** Storage optimizations, later. |
There was a problem hiding this comment.
please explain what are these dont talk about anything without defining it and explaining th econtext. what the hell is a content-addressed dedup or transcoding and why does it matter
There was a problem hiding this comment.
🤖 The AI agent says:
Defined now, each in one sentence in scope.md. Deduplication by content: store one copy when the same file is shared twice (a storage saving, not a correctness need). Transcoding: converting a file's format before delivery, for example turning a video into frames. Thumbnails: server-made preview images for large files. All three are deferred optimizations that do not block the core flow, which is why they are out of scope.
|
|
||
| ## The core idea: a reference on the wire, two objects behind it | ||
|
|
||
| The wire stops carrying base64. A shared file becomes a small **attachment reference** |
There was a problem hiding this comment.
you need to precise here. what do you mean by attachement reference. you are providing a json without any context to where it is used. is it in vercel? which part of the system. please explain the flow from a to z when you explain something and explain the boundaries of the system and the responsabilities of each tihng. explain high level but also mentioning the entities and primitives we have, this is in no a way an explanation a staff engineer would use.
There was a problem hiding this comment.
🤖 The AI agent says:
Rewritten, and the shape itself changed during review. The reference is no longer a raw {mount_id, path, ...}: the API now issues an opaque attachment resource on upload (attachment_id plus server-verified filename, media type, and size). The front end sends the attachment_id in the message parts in place of today's base64 file part; it rides the /run wire as a content block; it is what records persist; and the runner resolves it through the API, which checks the attachment belongs to the session. Raw storage coordinates would have let a client forge references to other sessions' files and made client-supplied metadata authoritative.
design.md now walks the flow end to end with every system named at each step (front end, API, object store, runner, sandbox, harness, model) and what each is responsible for; research.md section 2 has the responsibility split per flow; decision D10 records the shape choice.
|
|
||
| The wire stops carrying base64. A shared file becomes a small **attachment reference** | ||
| `{ mount_id, path, mediaType, filename, size }`. Behind that reference live **two objects with | ||
| different rights**, which is what lets "the user always finds it" and "the agent can do anything |
There was a problem hiding this comment.
what does the different right even mean
There was a problem hiding this comment.
🤖 The AI agent says:
Removed the phrase. The concrete difference: the original cannot be changed by the agent (it is outside the sandbox and only enters through a create-only API route), and the working copy can be changed freely (it sits in the working directory, which the agent fully controls). design.md decision D4 states it this way.
| different rights**, which is what lets "the user always finds it" and "the agent can do anything | ||
| to it" both be true at once: | ||
|
|
||
| - **The original** — *immutable, agent-unreachable.* Uploaded once to a session-scoped |
There was a problem hiding this comment.
why did we decide to have it immutable and agent unreachable etc.. what is the reason / constraint behind this design?
There was a problem hiding this comment.
🤖 The AI agent says:
The reason is the findability goal: a person must always be able to return to exactly what they shared. The agent's working directory is last-writer-wins, and its own README says concurrent runs overwrite freely. An agent asked to "clean up the workspace" can delete or overwrite anything there. Keeping the original where the agent cannot reach it is the only way the guarantee holds no matter what the agent does. design.md decision D3 spells this out with the concrete failure scenario.
| orig["<b>Original</b><br/>attachments mount<br/>immutable · agent-unreachable<br/><i>mounts/<proj>/<attach_mount>/…</i>"] | ||
| end | ||
| subgraph sandbox["Agent sandbox"] | ||
| cwd["<b>Working copy</b><br/>cwd/attachments/…<br/>mutable · last-writer-wins"] |
There was a problem hiding this comment.
if this is under project mount, how come is it unreachable. we want to make project mount optionall mountable to agents and accessible! and why is it in the proejct mount if it is session scoped?
There was a problem hiding this comment.
🤖 The AI agent says:
This was a misleading label in the diagram, and it is now fixed head-on. The storage key layout is mounts/<project_id>/<mount_id>/<path>, and <project_id> is only the tenant partition of the S3 key. It does not mean a "project mount". No project-scoped mount kind exists today; every mount in practice is scoped to a session or an agent. The attachments mount is session-scoped, and it is unreachable simply because it is never added to the sandbox's visible mount set; the project id in its key has nothing to do with that.
research.md section 2 states this explicitly, and the design.md prose around the diagram explains the label. A future project-level drive that is optionally mountable into agents is a separate feature that does not exist yet and is untouched by this design.
|
|
||
| ref -.addresses.-> orig | ||
| orig ==>|runner reads → base64| acp["ACP image / audio / document block<br/>→ model perceives"] | ||
| orig ==>|runner materializes copy| cwd |
There was a problem hiding this comment.
what does materializes copy here mean
There was a problem hiding this comment.
🤖 The AI agent says:
Defined in the README glossary and used consistently: to materialize means the runner writes a copy of the stored original into the agent's working directory (at cwd/attachments/<attachment_id>/<filename>) so the agent's tools can open it. Re-materialization only restores a missing copy; it never overwrites one the agent has edited.
| ref["Attachment reference<br/>{mount_id, path, mediaType, …}<br/><i>on the wire + in history + in traces</i>"] | ||
|
|
||
| ref -.addresses.-> orig | ||
| orig ==>|runner reads → base64| acp["ACP image / audio / document block<br/>→ model perceives"] |
There was a problem hiding this comment.
so the model perceives the data without needing to read it? is it different from the model perceiving the data after reading it?
There was a problem hiding this comment.
🤖 The AI agent says:
Yes, they are different, and the difference drives the design. Perceiving without reading: the bytes are delivered inside the message content and the model's own encoder sees them directly (native perception). Perceiving after reading: the agent runs a Read tool on a file and the model sees whatever text the tool returns. For images the second path is unreliable (Claude Code's Read tool does not dependably return an image as vision input), and for audio it does not exist at all. So native inline delivery is the dependable way to make the model perceive a file, and the disk copy serves the separate purpose of tool use. research.md sections 4 and 5.
| orig ==>|runner reads → base64| acp["ACP image / audio / document block<br/>→ model perceives"] | ||
| orig ==>|runner materializes copy| cwd | ||
| cwd ==>|Read / Bash / edit| tools["agent tools operate"] | ||
| orig ==>|render + download| ui["conversation + Shared-files drawer"] |
There was a problem hiding this comment.
what does render + download means
There was a problem hiding this comment.
🤖 The AI agent says:
Plain version: render means the conversation shows the file inline where it was sent (the front end resolves the reference to a download URL and displays it); download means the person gets the exact original bytes back from the download endpoint. Both are backed by the unchanging original. design.md, the findability section.
|
|
||
| Why the original must live **outside** the agent-writable tree: `cwd` is last-writer-wins | ||
| ([context.md](context.md)) — the agent can delete anything in it. If the shared file lived in | ||
| `cwd`, "always findable" would be at the agent's mercy. Because geesefs mounts a whole mount's |
There was a problem hiding this comment.
I dont understand this argument. what does it mean to be in the agent mercy. this whole paragraph is very badly written. please rewrite the argument with examples of what would fail if it were saved to the session mount.
I want to understand a scenario where something would break in case the file would be uplaoded in the session folder say under _uploads
There was a problem hiding this comment.
I would also want to understand if going through the simple implementaiotn would allow us in the future to go to th emore complex implemantoin progressively. or is it a different architecture.
There was a problem hiding this comment.
🤖 The AI agent says:
Rewritten around exactly your scenario. If the original lived at cwd/_uploads/report.pdf: you share report.pdf, later ask the agent to "clean up the workspace", and it runs rm -rf _uploads, or it overwrites report.pdf with a derived version while doing the work. The original you shared is gone, and "you can always find what you shared" is broken. That is the only thing that breaks in that world: perception and tool use would still work fine from _uploads. Which is exactly the point the design now makes honestly: the findability guarantee, and nothing else, is what forces the original out of the agent-writable tree. design.md decision D3, including the note that on lifecycle grounds alone a cwd subfolder would have been the simplest choice.
There was a problem hiding this comment.
🤖 The AI agent says:
It grows, with one honest caveat. Both versions change the same single runner call: the inline-only version (read bytes off the wire, build native content blocks) and the full version (resolve a reference, also materialize a working copy) replace the same hard-coded session.prompt line, so that boundary is preserved. The caveat: moving from inline bytes to durable references later is still substantial work around that call (front-end persistence, wire types, the record schema, storage ownership, authorization, cleanup). So it is the same architecture with a real migration, not a rebuild. design.md decision D9 states it this way, and the plan packages storage plus image delivery as one release partly for this reason.
|
|
||
| --- | ||
|
|
||
| ## Interaction 1 — end-to-end happy path (image + text, warm turn) |
There was a problem hiding this comment.
please write text explaining all these interactions and mermaid. again with context, flow, argumentation, why, how.
pelase add other options to undersand tradeoffs.
no need to remove the mermain just add
There was a problem hiding this comment.
🤖 The AI agent says:
Done. Every diagram in design.md now has prose before it: the context, the flow step by step, and why it is shaped that way. The major choices are written as options, what breaks under each, the decision, and the reason, and decisions.md holds the complete compact log D1 through D11 (two small ones, D2 and D8, are stated inline where they arise rather than as their own sections). The alternatives you asked for are in: file-on-disk versus inline delivery (D1), a cwd/_uploads folder versus the agent-files mount versus a dedicated mount (D3), one copy versus two (D4), raw storage coordinates versus an opaque attachment id (D10), and perception-only versus durable-input for the first release (D11, left open for you). All 8 diagrams are kept and re-validated with mermaid-cli.
…ter review Restructure into the plan-feature layout (context, research, design, plan, scope, decisions, status; proposal.md absorbed), answer every review comment, verify the pinned adapter behavior (images native on both; audio and blob documents undelivered today), switch the wire reference to a server-issued attachment id, fix the upload flow to create-only API routes, make capability a three-layer intersection, and repackage the plan into one user-visible release with security and limits included. D11 (what the first release promises) is left open for the product owner. Claude-Session: https://claude.ai/code/session_01Jg92VZvmkMnm5VQXprTNEH
|
🤖 The AI agent says: I reworked this design workspace to address every review comment. Each of your inline comments has a reply in its thread explaining the answer and what changed; this comment is the overview. One decision is left open for you (D11, at the end). The new structureThe workspace now follows the plan-feature layout.
What changed in the design itself, beyond reorganizationWorking through your comments forced real changes, not just rewriting.
The decision that needs you: D11What does attaching a file promise in the first release? The engineering design supports either answer; the plan is written for the recommended one and says explicitly how Stage 1 shrinks if you choose the other.
The recommendation is Option B, because findability is one of your stated product goals, the three payload problems are real today, and Option A ships a UI whose promises change under the user later. It is your call; decisions.md D11 is marked open until you decide. |
mmabrouk
left a comment
There was a problem hiding this comment.
🤖 The AI agent says:
These inline notes mark where each of your review questions is now answered in the reworked documents. The overview of all changes is in the summary comment.
| (`agent-mount.ts:23`). In other words, the agent can overwrite or delete anything in `cwd`. That | ||
| fact drives a key design choice (see [design.md](design.md), the decision on where originals live). | ||
|
|
||
| ### Every mounts endpoint, who may call it, and what it does |
There was a problem hiding this comment.
🤖 The AI agent says:
This endpoint table, and the write-path and responsibility sections after it, answer your research request in this comment: every mounts route with its permission and arguments, the two byte paths (API versus signed credentials), the drawer flows, and who owns each step. It also resolves the project-mount confusion from this comment: the project id is only the tenant partition of the storage key, and no project-scoped mount kind exists.
|
|
||
| --- | ||
|
|
||
| ## 3. Kinds of files and how models handle them |
There was a problem hiding this comment.
🤖 The AI agent says:
This section is the general modality taxonomy you asked for in this comment (images, audio, video, PDFs, Excel, zips, CAD, folders, large files) and the native-versus-just-files distinction from this one: native means the model's own encoder must receive the bytes through message content; everything else is a file for the agent's tools. Provider limits, with sources, are at the end of the section.
| `capabilities.ts:86`). So the runner receives the capability values but never uses them. This | ||
| project makes the runner enforce them. More on the history in section 9. | ||
|
|
||
| ### What the pinned adapters actually do |
There was a problem hiding this comment.
🤖 The AI agent says:
This subsection answers your ACP question in this comment with the actual code of the two adapters we pin, not with the spec alone. The load-bearing findings: images arrive at the model natively through both adapters; neither supports native audio; the Claude adapter drops blob resources entirely, so a PDF sent as an embedded resource never reaches Claude today. This settles the old open question about document delivery and is why the audio-and-documents stage is blocked on adapter work.
|
|
||
| --- | ||
|
|
||
| ## 7. Records and replay |
There was a problem hiding this comment.
🤖 The AI agent says:
This section captures the direction from your comment: the full-history resend is pre-sessions behavior, and the runner will own rebuilding context from records. The issue you asked for in the follow-up is filed as #5443. One fact that feeds the plan: records are text-only today, so the record-schema extension is part of the first release.
|
|
||
| --- | ||
|
|
||
| ## Decision D3: where the original lives |
There was a problem hiding this comment.
🤖 The AI agent says:
This decision answers your why-immutable question from this comment and contains the concrete cwd/_uploads failure scenario you asked for in this one: only the findability guarantee breaks if the original lives in the working directory, and that is exactly why it is the one thing that forces the original out. The note at the end concedes that on lifecycle grounds alone a cwd subfolder would have been the simplest choice.
|
|
||
| --- | ||
|
|
||
| ## Decision D6: what does attaching a file promise, and what happens on an unsupported kind |
There was a problem hiding this comment.
🤖 The AI agent says:
This decision is where your point from this comment reshaped the design: attach always puts the file in the workspace (tools work regardless of the model), native delivery happens when the three-layer capability intersection allows it, an unsupported kind becomes a workspace-only attachment with a visible notice, and the runner fails a turn only on a contract violation. No silent drops, and no blocking a non-vision model from tool use.
|
|
||
| --- | ||
|
|
||
| ## Decision D11 (open): what does attaching a file promise in the first release? |
There was a problem hiding this comment.
🤖 The AI agent says:
This is the one decision that needs you: what the first release promises (image perception only, or durable agent input). The options and the recommendation are here and in decisions.md D11; the plan states how Stage 1 shrinks if you pick the smaller option.
|
|
||
| --- | ||
|
|
||
| ## Decision D10: what the reference contains |
There was a problem hiding this comment.
🤖 The AI agent says:
This decision answers your what-is-an-attachment-reference question from this comment, and the shape changed while answering it: instead of raw storage coordinates, the API now issues an opaque attachment_id with server-verified metadata, which prevents forged references to other sessions' files and stops client-supplied media types from being authoritative. The end-to-end flow with every system named is in the successful-flow section above.
|
|
||
| This stage can be skipped if Stage 1 lands quickly. | ||
|
|
||
| ## Stage 1: the first user-visible release (images perceived, files travel as references) |
There was a problem hiding this comment.
🤖 The AI agent says:
Stage 1 folds in the tracing change you asked to see planned in this comment (the trace carries a reference once messages do), plus the record-schema extension and the minimum security and limits work. The packaging is deliberate: storage and model perception ship as one visible release so the UI never accepts a file the model ignores.
| | D6 | What happens when the model cannot perceive an attached kind? | refuse the attachment; drop it silently; fail the turn; attach it as a workspace-only file with a visible notice | Attach it as a workspace-only file with a visible notice; the runner fails the turn only on a contract violation (asked to deliver a native block the harness cannot accept) | Attach means two things: show it to the model, and put it in the workspace. The workspace half always works, so refusing or failing would block legitimate tool use. Silent dropping is the current trap. A visible notice keeps it honest. | | ||
| | D7 | How is the original kept immutable? | keep it out of the sandbox; also write only through a create-only upload route and never sign the mount; also a read-only credential scope | Keep it out of the sandbox, and write only through the create-only upload route with no signed credentials for the mount; read-only scope is a hardening follow-up | Keeping it out of the sandbox stops the agent, the main threat. Originals enter only through the create-only upload route, and the API refuses overwrite and delete of an original as an explicit check, because `write_file` overwrites silently today. No signed credentials are issued for the mount, because any signed credential is read-write today. A read-only scope is the strongest form and needs a signing change, so it waits. | | ||
| | D8 | Is audio in scope? | defer; include | Include as a product goal, but blocked on adapter work | Audio stays a goal, and there is no disk-read fallback for it, so it forces a real inline audio block. Neither pinned adapter (`claude-agent-acp`, `pi-acp`) advertises native audio today, so the audio stage cannot ship against the current pins and is blocked on adapter or harness work. | | ||
| | D9 | Can the inline-only version grow into the full one? | it is a different architecture; it grows cleanly with no rework; the model-facing seam is preserved but the rest is system-wide | The model-facing seam is preserved, but adopting durable references later is still a system-wide change | Shipping the inline-only version first avoids a rewrite of the one prompt-builder seam. It does not avoid the rest: front-end persistence, API storage ownership, SDK wire types, runner resolution, the record schema, authorization, and cleanup. So the benefit is narrow and honest, not a free migration. | |
There was a problem hiding this comment.
🤖 The AI agent says:
This row answers your progressive-path question from this comment, with the honest version: the one runner call is preserved between the inline-only and full versions, but adopting durable references later is still a system-wide change. Same architecture, real migration.
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 7
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: 690bf567-e47b-4a7a-a577-87fd30afe36d
📒 Files selected for processing (8)
docs/design/agent-workflows/projects/agent-multi-modality/README.mddocs/design/agent-workflows/projects/agent-multi-modality/context.mddocs/design/agent-workflows/projects/agent-multi-modality/decisions.mddocs/design/agent-workflows/projects/agent-multi-modality/design.mddocs/design/agent-workflows/projects/agent-multi-modality/plan.mddocs/design/agent-workflows/projects/agent-multi-modality/research.mddocs/design/agent-workflows/projects/agent-multi-modality/scope.mddocs/design/agent-workflows/projects/agent-multi-modality/status.md
| **Decision.** Option C. A dedicated attachments mount, scoped to the session, that is created with | ||
| the existing `get_or_create_session_mount(session_id, name="attachments")` machinery but is | ||
| deliberately never added to the set of mounts made visible in the sandbox. The runner reads it | ||
| out-of-band with the signed credentials it already obtains, using a plain object GET rather than a | ||
| folder mount (this is possible; see [research.md](research.md), section 2). |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift
Align D3 with the immutability decision in D7.
This section says the runner reads the attachments mount with signed credentials, but D7 requires that no signed credentials ever be issued for that mount because current credentials are read-write. Stage 1 and the sequence diagram instead use the session-bound API download route. Choose one contract; the safe first-release path is API download, with read-only credentials deferred to Stage 3.
| Referenced --> Materialized: runner always writes the working copy | ||
| Referenced --> Perceived: native delivery when the capability intersection allows | ||
| Referenced --> WorkspaceOnly: intersection does not allow, delivered as a workspace-only attachment with a notice |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Clarify that re-materialization occurs only when the copy is missing.
“Runner always writes the working copy” conflicts with the explicit rule at Lines [218-221] that edited copies must never be overwritten. Change this transition to “materializes the working copy if missing” so implementation cannot destroy agent edits.
| - The create-only upload route: get-or-create the attachments mount, validate the file (size, count, | ||
| media type verified server-side by inspecting the file bytes, not by trusting the client), write the bytes, and | ||
| return `{attachment_id, filename, media_type, size}`. Refuse an overwrite or a delete of an | ||
| existing attachment original, because `write_file` overwrites silently today |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
Define upload idempotency and atomic commit semantics.
The tests at Lines [298-300] require retries to avoid duplicates and partial originals, but “create-only” and mount get-or-create do not provide that guarantee. Specify an idempotency key/client upload token and a temporary-write-then-finalize flow (or equivalent) before relying on this behavior.
| **Tracing.** Because the history now carries a reference, the Python span no longer holds base64. | ||
| Confirm the trace shows the reference, not the bytes (the Python agent span keeps `messages` on | ||
| purpose, and a reference there is small). |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | 🏗️ Heavy lift
Prevent resolved attachment bytes from entering traces.
Replacing history bytes with references only protects traces that record the logical message. The runner later resolves each reference into base64 ACP content, so define the tracing boundary explicitly: record references/metadata, redact resolved data/blob fields, and add a regression test that traces contain no attachment bytes.
🧰 Tools
🪛 LanguageTool
[style] ~113-~113: Try using a descriptive adverb here.
Context: ...(the Python agent span keeps messages on purpose, and a reference there is small). **SD...
(ON_PURPOSE_DELIBERATELY)
| - **Resending the whole conversation every turn.** The front end resends the full message history | ||
| on each turn (`agentRequest.ts`, the `history` array). A five megabyte image becomes about | ||
| seven megabytes of base64 and is re-uploaded on every later turn. The cost grows with the | ||
| square of the conversation length. |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Qualify the quadratic-cost claim.
The cumulative history payload can be quadratic in turn count in the worst case, but repeatedly resending one fixed attachment is linear in the number of later turns. Replace the unconditional statement with “can grow quadratically with conversation length” or state the assumptions.
| - **Images** (PNG, JPEG, GIF, WebP). Perceived by Claude, the GPT-4o and GPT-5 family, and Gemini. | ||
| - **Audio** (for example WAV, MP3). Perceived natively by some models (Gemini, and the GPT audio | ||
| models). There is no reliable way to make a model "hear" audio by having the agent read the file, | ||
| so audio must be delivered as a native audio input. | ||
| - **PDF and other documents as a document input.** Claude accepts a PDF as a document content block | ||
| and reads both its text and its page images. This is a native path distinct from plain text. | ||
| - **Video.** Gemini perceives video natively. Claude and the current GPT chat models do not. Video | ||
| is out of scope for this project (see [scope.md](scope.md)). |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
sed -n '220,280p' docs/design/agent-workflows/projects/agent-multi-modality/research.mdRepository: Agenta-AI/agenta
Length of output: 4018
🏁 Script executed:
nl -ba docs/design/agent-workflows/projects/agent-multi-modality/research.md | sed -n '244,266p'
rg -n "GPT-4o|GPT-5|video|Claude accepts|Gemini perceives" docs/design/agent-workflows/projects/agent-multi-modality/research.mdRepository: Agenta-AI/agenta
Length of output: 975
🏁 Script executed:
sed -n '300,325p' docs/design/agent-workflows/projects/agent-multi-modality/research.md
sed -n '470,635p' docs/design/agent-workflows/projects/agent-multi-modality/research.mdRepository: Agenta-AI/agenta
Length of output: 11667
Add direct sources for the model-capability claims here. The Anthropic and Gemini links later in the section cover the limit notes, not the GPT-4o/GPT-5 or video-support statements. Either cite vendor docs here or narrow the list to sourced examples.
| [#5443, "(feat) Runner should rebuild session context from records when a session cannot be resumed warm"](https://github.com/Agenta-AI/agenta/issues/5443). | ||
| The related warm case is | ||
| [#5384, "(feat) Warm sessions should hold client tool calls open instead of replaying the turn"](https://github.com/Agenta-AI/agenta/issues/5384): | ||
| #5384 avoids a replay when the session is still alive, while #5443 covers rebuilding from records |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Fix the malformed issue reference.
#5384 has no space after the hash and triggers MD018. Write Issue #5384`` or otherwise escape the reference so Markdown lint passes.
🧰 Tools
🪛 markdownlint-cli2 (0.23.0)
[warning] 555-555: No space after hash on atx style heading
(MD018, no-missing-space-atx)
Source: Linters/SAST tools
Context
Agent workflows are text-only at the model. A person can attach an image or a document in the agent chat, the bytes travel intact from the browser through the API and SDK to the runner, and the runner then drops everything except text at the single call that hands the turn to the harness (
run-turn.ts:742, hard-coded to one text block). The model never sees the file, an image-only message is rejected outright, and nothing warns the person. The prompt (non-agent) playground handles files correctly, so agents are behind the rest of the product and the gap is invisible until the agent ignores your file.This PR contains the design workspace for closing that gap, under
docs/design/agent-workflows/projects/agent-multi-modality/. It is documentation only: no production code changes.The design in one paragraph
The wire stops carrying file bytes. The front end uploads a file once through a create-only API route; the API stores it in a session-scoped attachments mount (our existing S3-backed mounts system) that is never exposed inside the sandbox, and returns an opaque, server-issued attachment id. Messages, records, and traces carry that id instead of base64. Behind the id sit two objects: an immutable original (what the model reads, what renders inline, what download returns) and a working copy the runner writes into the agent's working directory so tools can open or edit it. At prompt time the runner resolves the id, builds native ACP content blocks (the protocol requires inline bytes for the model to perceive a file), and replaces the hard-coded text block. Whether a modality reaches the model is the intersection of three layers (protocol transport, adapter fidelity, model modalities); a file the model cannot perceive is still attached as a workspace file with a visible notice, never silently dropped.
How to review
Read in this order:
claude-agent-acp0.58.1,pi-acp0.0.29) actually do with each block type, how Zed and opencode handle attachments, the records system, and the capability-flag history.The one decision this PR needs
D11 (open): what does attaching a file promise in the first release? Image perception only (smallest fix, storage migration later), or durable agent input (immutable original, working copy, references in records; recommended). The options, tradeoffs, and recommendation are in decisions.md and in the summary comment below. Everything else in the design is decided, with the reasoning recorded.
Verified facts worth knowing before reviewing
https://claude.ai/code/session_01Jg92VZvmkMnm5VQXprTNEH