Feature request: host-managed embeddings API for Anna Apps and Executa Tools

Hi Anna team,

I am working on an Anna App that adapts a research workflow into the Anna runtime. The current flow is roughly:

  1. The user enters a research question.
  2. The app runs web search and receives candidate sources.
  3. The tool/backend extracts or summarizes source content.
  4. The app splits content into text chunks.
  5. The app ranks or filters those chunks against the research question.
  6. The selected context is passed to `anna.llm.complete(…)` to generate the final research report.

For step 5, a standard implementation is embedding-based retrieval: embed the query, embed the text chunks, then compute similarity locally and keep the most relevant chunks. Today, that requires each app/tool to bring its own embedding provider, API key, model routing, billing, quota handling, and fallback behavior.

Anna already provides host-managed LLM access through `anna.llm.complete(…)` in the app iframe and `sampling/createMessage` from Executa tools. It would be very useful to have a similar host-managed embeddings capability so Anna Apps can build retrieval, RAG, semantic deduplication, clustering, and context selection without requiring separate third-party embedding credentials.

Requested capability

Please consider adding an embeddings API owned by the Anna host. The app/tool would send plain text inputs and receive vectors. Similarity ranking, threshold filtering, vector cache, and local vector-store integration can remain app-side.

Suggested App iframe API

One possible shape:

```ts
const response = await anna.llm.embed({
input: [
“original research question”,
“first source text chunk”,
“second source text chunk”
],
purpose: “retrieval”,
metadata: {
app: “anna-researcher”,
stage: “context_selection”
}
});
```

Suggested response:

```ts
{
model: “anna-managed-embedding-model”,
dimensions: 1536,
vectors: [
{ index: 0, embedding: [0.0123, -0.044, 0.281] },
{ index: 1, embedding: [0.0198, -0.031, 0.244] }
],
usage: {
input_tokens: 1234
}
}
```

An alternative name such as `anna.embeddings.create(…)` would also work. The important parts are batch input, stable vector output, usage metadata, and host-managed model selection.

Suggested Executa reverse RPC API

For tools, the equivalent could mirror the existing sampling pattern:

```json
{
“jsonrpc”: “2.0”,
“id”: “plugin-generated-id”,
“method”: “embeddings/create”,
“params”: {
“input”: [
“research question”,
“document chunk 1”,
“document chunk 2”
],
“purpose”: “retrieval”,
“metadata”: {
“executa_invoke_id”: “8f1c…”,
“tool”: “select_context”,
“stage”: “embedding”
}
}
}
```

Suggested result:

```json
{
“jsonrpc”: “2.0”,
“id”: “plugin-generated-id”,
“result”: {
“model”: “anna-managed-embedding-model”,
“dimensions”: 1536,
“vectors”: [
{ “index”: 0, “embedding”: [0.0123, -0.044, 0.281] },
{ “index”: 1, “embedding”: [0.0198, -0.031, 0.244] }
],
“usage”: {
“input_tokens”: 1234
}
}
}
```

The manifest capability could be something like:

```json
{
“host_capabilities”: [“llm.embed”]
}
```

or:

```json
{
“host_capabilities”: [“embeddings.create”]
}
```

For Executa initialization, a tool could advertise:

```json
{
“capabilities”: {
“embeddings”: {}
}
}
```

Integration details that would make this easy to adopt

  1. Support batch input: `input` should accept `string | string[]`.
  2. Return `dimensions` so apps can validate caches and vector indexes.
  3. Document whether vectors are normalized and whether cosine similarity, dot product, or another metric is recommended.
  4. Allow Anna to choose a default embedding model when the app does not pass `model`.
  5. Optionally support model preferences later, for example `{ task: “retrieval”, quality: “balanced” }`, without requiring apps to bind to a provider-specific model name.
  6. Include usage metadata for quota and debugging.
  7. Support request metadata such as `executa_invoke_id`, `tool`, and `stage` for auditing.
  8. Clearly document limits: maximum items per batch, maximum tokens per item, maximum tokens per request, and whether over-limit input fails or is truncated.
  9. Provide stable error codes for capability not granted, quota exceeded, input too large, model unavailable, invalid input, and transient provider error.
  10. Keep credentials fully host-managed. The app/tool should never receive or manage the underlying embedding provider key.

Why this matters for Anna Apps

For a research app, the desired flow is:

```text
web search results
→ text chunks
→ Anna embeddings
→ local similarity ranking/filtering
→ selected context
→ anna.llm.complete report generation
```

Only the embedding generation needs platform support. The app can do the rest locally. This would make it much easier to build high-quality research tools, knowledge-base search, document Q&A, semantic deduplication, and context compression while keeping model routing, billing, quota, and user authorization inside Anna.

Is an embeddings API already on the roadmap? If yes, I would appreciate guidance on the preferred API shape so app developers can design against it early.

Hi HappyLight,

Great news! :tada: The anna-app development kit now supports embedding models in v0.1.19!

This means you can now build your research workflow with host-managed embeddings for retrieval, RAG, and context selection without needing separate third-party embedding credentials.

Thank you so much for your detailed feature request and thoughtful API suggestions - they were incredibly helpful in shaping this implementation! :blush:

Wed love to hear more of your suggestions as you start using the new embedding capabilities! :light_bulb: