Hi Anna team,
I am working on an Anna App that adapts a research workflow into the Anna runtime. The current flow is roughly:
- The user enters a research question.
- The app runs web search and receives candidate sources.
- The tool/backend extracts or summarizes source content.
- The app splits content into text chunks.
- The app ranks or filters those chunks against the research question.
- The selected context is passed to `anna.llm.complete(…)` to generate the final research report.
For step 5, a standard implementation is embedding-based retrieval: embed the query, embed the text chunks, then compute similarity locally and keep the most relevant chunks. Today, that requires each app/tool to bring its own embedding provider, API key, model routing, billing, quota handling, and fallback behavior.
Anna already provides host-managed LLM access through `anna.llm.complete(…)` in the app iframe and `sampling/createMessage` from Executa tools. It would be very useful to have a similar host-managed embeddings capability so Anna Apps can build retrieval, RAG, semantic deduplication, clustering, and context selection without requiring separate third-party embedding credentials.
Requested capability
Please consider adding an embeddings API owned by the Anna host. The app/tool would send plain text inputs and receive vectors. Similarity ranking, threshold filtering, vector cache, and local vector-store integration can remain app-side.
Suggested App iframe API
One possible shape:
```ts
const response = await anna.llm.embed({
input: [
“original research question”,
“first source text chunk”,
“second source text chunk”
],
purpose: “retrieval”,
metadata: {
app: “anna-researcher”,
stage: “context_selection”
}
});
```
Suggested response:
```ts
{
model: “anna-managed-embedding-model”,
dimensions: 1536,
vectors: [
{ index: 0, embedding: [0.0123, -0.044, 0.281] },
{ index: 1, embedding: [0.0198, -0.031, 0.244] }
],
usage: {
input_tokens: 1234
}
}
```
An alternative name such as `anna.embeddings.create(…)` would also work. The important parts are batch input, stable vector output, usage metadata, and host-managed model selection.
Suggested Executa reverse RPC API
For tools, the equivalent could mirror the existing sampling pattern:
```json
{
“jsonrpc”: “2.0”,
“id”: “plugin-generated-id”,
“method”: “embeddings/create”,
“params”: {
“input”: [
“research question”,
“document chunk 1”,
“document chunk 2”
],
“purpose”: “retrieval”,
“metadata”: {
“executa_invoke_id”: “8f1c…”,
“tool”: “select_context”,
“stage”: “embedding”
}
}
}
```
Suggested result:
```json
{
“jsonrpc”: “2.0”,
“id”: “plugin-generated-id”,
“result”: {
“model”: “anna-managed-embedding-model”,
“dimensions”: 1536,
“vectors”: [
{ “index”: 0, “embedding”: [0.0123, -0.044, 0.281] },
{ “index”: 1, “embedding”: [0.0198, -0.031, 0.244] }
],
“usage”: {
“input_tokens”: 1234
}
}
}
```
The manifest capability could be something like:
```json
{
“host_capabilities”: [“llm.embed”]
}
```
or:
```json
{
“host_capabilities”: [“embeddings.create”]
}
```
For Executa initialization, a tool could advertise:
```json
{
“capabilities”: {
“embeddings”: {}
}
}
```
Integration details that would make this easy to adopt
- Support batch input: `input` should accept `string | string[]`.
- Return `dimensions` so apps can validate caches and vector indexes.
- Document whether vectors are normalized and whether cosine similarity, dot product, or another metric is recommended.
- Allow Anna to choose a default embedding model when the app does not pass `model`.
- Optionally support model preferences later, for example `{ task: “retrieval”, quality: “balanced” }`, without requiring apps to bind to a provider-specific model name.
- Include usage metadata for quota and debugging.
- Support request metadata such as `executa_invoke_id`, `tool`, and `stage` for auditing.
- Clearly document limits: maximum items per batch, maximum tokens per item, maximum tokens per request, and whether over-limit input fails or is truncated.
- Provide stable error codes for capability not granted, quota exceeded, input too large, model unavailable, invalid input, and transient provider error.
- Keep credentials fully host-managed. The app/tool should never receive or manage the underlying embedding provider key.
Why this matters for Anna Apps
For a research app, the desired flow is:
```text
web search results
→ text chunks
→ Anna embeddings
→ local similarity ranking/filtering
→ selected context
→ anna.llm.complete report generation
```
Only the embedding generation needs platform support. The app can do the rest locally. This would make it much easier to build high-quality research tools, knowledge-base search, document Q&A, semantic deduplication, and context compression while keeping model routing, billing, quota, and user authorization inside Anna.
Is an embeddings API already on the roadmap? If yes, I would appreciate guidance on the preferred API shape so app developers can design against it early.