Proposal: Provide first-class attachment parsing tools for Agent Sessions (Excel / DOCX / PPTX)

Background

I am developing an Anna App for generating PowerPoint presentations (this repository contains the current project). One of our core requirements is to let users upload attachments and have the Agent use their contents in the subsequent authoring workflow. For example:

  • A user may upload a PPTX and ask the App to use it as a reference for visual style or slide organization.
  • A user may upload an Excel workbook and ask the App to generate a presentation from its data, tables, or analysis.
  • A user may upload a DOCX document and ask the App to extract its content, tables, and structure before generating a presentation.

For these workflows, Agent Sessions need reliable and reusable attachment-reading and understanding capabilities.

Current behavior observed

I created and ran a raw Agent Session in examples/test-session-app and sent the following request:

Please read and understand the contents of /Users/leyouming/anna-workspace/tmp/测试excel.xlsx.

The Session was able to access the file, but it did not use a dedicated Excel parsing capability. Instead, it invoked exec_run, installed and used Python, pandas, and openpyxl, and ran code similar to:

df = pd.read_excel(path, sheet_name=None)
print({sheet: data.head().to_dict() for sheet, data in df.items()})

This creates several concrete problems:

  1. The result is incomplete: head() returns only the first five rows. For a larger workbook, the Agent may silently inspect only a subset of the data and then produce conclusions based on that subset.
  2. Types and semantics can be lost: Excel dates were returned as serial numbers (for example, 46233). The Session eventually interpreted that value as “around August 2025”, while the workbook actually displays dates in July 2026.
  3. Results depend on ad hoc code written by the Agent: Different models may choose different libraries, read options, and truncation strategies, producing different results for the same file.
  4. The execution environment has extra requirements: This approach depends on Python, pandas, openpyxl, and the ability to install or download packages. It may fail in a restricted, offline, or Python-less environment.
  5. There is no auditable source location: The output is a temporary printed representation without stable references such as worksheet/cell ranges, paragraph or table indexes, page numbers, or slide numbers. This makes it difficult to trace facts back to the source attachment during presentation generation.

In this test, the original Excel workbook contained one worksheet and 17 data records. The Session inspected only the first five records. It therefore achieved a partial read, but it did not fully understand the workbook.

Requested platform capability

I would like Anna to provide platform-maintained attachment parsing tools through the existing Executa/tool model. The goal is not for every App to implement its own parsing protocol, libraries, and runtime setup, but for Anna to provide stable, reusable official tools that Sessions can call when they need to process common attachment formats.

The tools should return structured and traceable results where possible, including:

  • File metadata: filename, MIME type, size, page count, or worksheet count;
  • Structured content: text, tables, lists, formulas, charts, and images;
  • Type information: dates, numbers, percentages, currencies, empty values, and hyperlinks;
  • Stable source locations, such as Sheet1!A1:D18 for Excel, paragraph/table indexes for DOCX, and slide numbers for PPTX;
  • Large-file handling information: whether the result was truncated, why it was truncated, and how to paginate or continue reading;
  • Optional rendered output, such as page PNGs, worksheet previews, or slide images, for visual inspection by multimodal models;
  • Artifact IDs, URLs, or paths for large results instead of placing large byte payloads directly into the model context.

Recommended initial formats and capabilities

Excel (.xlsx / .xls)

  • Worksheet list, hidden state, used ranges, and table objects;
  • Displayed values, raw values, data types, and formulas for cells;
  • Correct conversion of Excel date serials;
  • Merged cells, filters, frozen panes, conditional formatting, and other structures that affect interpretation;
  • Metadata for charts, images, comments, data validation, and named ranges;
  • Pivot tables or other common summary structures, if they can be supported reliably;
  • Optional worksheet rendering previews;
  • Pagination by worksheet or cell range for large workbooks.

DOCX (.docx)

  • Heading hierarchy, paragraphs, lists, tables, hyperlinks, and footnotes;
  • Headers, footers, tables of contents, page breaks, and common style information;
  • Images, text boxes, and their positional relationships;
  • Revisions, comments, and document properties, clearly separated from the main document content where applicable;
  • Page-level rendering for preserving layout semantics;
  • Source locations at paragraph or table-cell level.

PPTX (.pptx)

  • Slide order, titles, body text, notes, tables, and charts;
  • Images, shapes, text boxes, SmartArt, and basic layout information;
  • Theme, master, font, color, and layout information useful for visual reference;
  • Speaker notes, hidden slides, and hyperlinks;
  • A rendered image for each slide to support visual inspection by the Agent;
  • Separate representations for “content extraction” and “visual reference”, so text or claims visible in an image are not automatically treated as verified factual evidence.

PDF (.pdf)

  • Page-level text, headings, lists, tables, and links;
  • Source locations by page, paragraph, or table region;
  • Page rendering for complex layouts and visual inspection of scanned documents;
  • OCR results for scanned PDFs, clearly marked with the fact that they are recognition results and may be uncertain.

Other common formats

If the platform already has suitable underlying libraries, support could be expanded over time to:

  • CSV / TSV: encoding, delimiter, column types, headers, and pagination for large files;
  • Markdown / HTML: heading hierarchy, lists, tables, links, code blocks, and image references;
  • ODT and other common office-document formats;
  • JSON / XML: structured expansion, field paths, and partial reads for large files.

Image attachments are not the main gap in this request. They can already be handled through analyze_image or the model’s multimodal capabilities. The missing platform-level capability is reliable processing for compound binary attachments such as Excel, DOCX, PPTX, and PDF files.

Expected benefits

Platform-level attachment parsing tools would:

  • Let different Anna Apps reuse the same reliable parsing capabilities;
  • Avoid requiring every App to install Python and maintain its own file-processing scripts;
  • Reduce omissions, type misinterpretations, and inconsistent results caused by ad hoc code generation;
  • Allow Agents to cite precise locations in the attachment, improving auditability of generated content;
  • Make large-file handling, timeouts, resource limits, and error recovery easier to manage consistently;
  • Provide shared infrastructure for presentations, reports, knowledge bases, document summarization, data analysis, and other file-based workflows.

Questions for the Anna team

  1. Does the Anna platform already have plans to provide first-class parsing tools for common non-text attachments such as Excel, DOCX, PPTX, and PDF?
  2. If such capabilities already exist, what is the recommended Session invocation pattern and tool ID?
  3. If they do not exist yet, could common attachment parsing be considered as a platform-level official Executa / Host Tool capability?

This capability would benefit far more than PPT-generation Apps; it would provide a common foundation for any Anna App that needs to process user-uploaded files. I would appreciate the team’s guidance on the API shape, supported formats, and prioritization.

Hi @HappyLight! :waving_hand:

Thank you for this outstanding write-up — the reproduction in examples/test-session-app was incredibly helpful. You identified every failure mode precisely: the ad-hoc exec_run + pandas path, the silent head() truncation (5 of 17 rows!), and the raw date serial 46233 being misread as 2025 instead of 2026. We fully agree this belongs in the platform, not in every App. :100:

Good news: this shipped in 1.1.0-beta.116 :tada:

Agent Sessions now have two first-class, server-side attachment parsing tools:

  • :bar_chart: sheet_read — Excel (.xlsx / .xls)
  • :page_facing_up: doc_read — DOCX / PPTX / PDF

They directly address each issue you raised:

  • No silent truncation, ever. Every response carries a truncated block with total_rows / total_units, and when a result is cut off, next_call contains the exact parameters to copy for the next page — the agent is instructed to keep paginating until it has read everything relevant. :repeat_button:
  • Correct types. Each cell returns {v, d, t} — raw value, displayed value, and type (date | number | percent | currency | text | bool | formula | empty | error). Dates are always ISO 8601, so no more serial-number guessing. :date:
  • Stable, citable locators. Sheet1!A2:D18 for Excel, para:N / table:M for DOCX, slide:N for PPTX, page:N for PDF — agents are prompted to attach locators when citing attachment content. :magnifying_glass_tilted_left:
  • Zero runtime requirements. Parsing runs server-side — no Python, pandas, or package installs needed in the execution environment.

How to use them (your question #2):

  1. Pass your attachment via attachments on agent.session.run as usual — the platform’s attachment hint now steers the model to these tools automatically, so most Apps need no changes at all. :sparkles:
  2. For Excel, the recommended pattern is two hops: call sheet_read with just the url for a workbook overview (sheet list, dimensions, header preview), then read specific sheets with pagination. doc_read similarly supports mode: "outline" for locating content in large documents before a full read.
  3. If your App uses an explicit allowed_tools list for sandbox sessions, add sheet_read and doc_read to it.
  4. For local development, update to CLI 0.1.45+ so the dev harness catalog includes both tools.

On the rest of your proposal — rendered previews (slide/page images), OCR for scanned PDFs, and additional formats like CSV/Markdown are not in this first release, but they’re very much on our radar (OCR is tracked in a separate thread). Scanned PDFs currently return an is_scanned_hint so the agent knows the text layer is empty rather than guessing.

Please give it a try with your PPT-generation workflow and let us know how it goes — feedback like this genuinely shapes the platform. Thanks again for raising it! :folded_hands::yellow_heart: