Running AI Models at the Edge with Workers AI: Summarization API, Vectorize RAG and D1
Key takeaways
How to call Workers AI models from a Cloudflare Worker, build a summarization API and a small retrieval-augmented chatbot on Vectorize, store conversations in D1, and avoid the mistakes that silently break retrieval or inflate the bill.
What Workers AI is
Workers AI lets a Cloudflare Worker run inference on models that Cloudflare hosts on its own GPUs: text generation (Llama, Mistral and others), embeddings, summarization, translation, speech recognition and image models. You do not provision GPUs or manage model weights. You add an AI binding to your Worker and call env.AI.run(modelId, input).
It pairs with other Cloudflare products that matter for AI features:
- Vectorize is a vector database for nearest-neighbor search over embeddings, which is the retrieval half of RAG.
- D1 is SQLite-based SQL storage for conversations, users and document text.
- R2 is S3-compatible object storage for source files.
- KV is an eventually consistent key-value store, useful for caching.
It helps to be precise about the “edge” claim. Your Worker code runs in a data center near the user, but the model runs wherever Cloudflare has GPU capacity for that model, which is not necessarily the same location. For a small model called from a Worker, total latency is often good. Still, do not design on the assumption that inference always happens in the user’s city. Measure from the regions your users are in.
The main trade-off compared with calling a large hosted model API is model choice. The catalog is mostly open-weight models in the small-to-mid size range. They are good at summarization, classification, extraction and short answers grounded in retrieved context. They are noticeably weaker than frontier models at long multi-step reasoning. Many production setups use both: Workers AI for cheap, high-volume tasks such as embeddings and classification, and an external API for the hard requests.
Setup and bindings
npm create cloudflare@latest my-ai-app
cd my-ai-app
npx wrangler login
The binding is what gives your code env.AI. It is the step most often missing from copied examples:
# wrangler.toml
name = "my-ai-app"
main = "src/index.ts"
compatibility_date = "2026-09-01"
[ai]
binding = "AI"
// src/index.ts
export interface Env {
AI: Ai;
}
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const result = await env.AI.run('@cf/meta/llama-3.1-8b-instruct-fp8', {
messages: [{ role: 'user', content: 'Explain HTTP caching in two sentences.' }],
max_tokens: 200,
});
return Response.json(result);
},
};
The Ai type comes from @cloudflare/workers-types (or from wrangler types, which generates an Env interface from your config). Using it instead of any gives you typed inputs per model id. For example, the compiler knows that @cf/facebook/bart-large-cnn takes input_text and returns summary.
wrangler dev runs your Worker locally, but AI calls still go to Cloudflare’s GPUs and use your account’s usage. There is no offline model emulator, so an integration test loop that calls the model on every run costs money.
A summarization endpoint
export interface Env {
AI: Ai;
}
export default {
async fetch(request: Request, env: Env): Promise<Response> {
if (request.method !== 'POST') {
return new Response('Method Not Allowed', { status: 405 });
}
let text: unknown;
try {
({ text } = await request.json<{ text?: string }>());
} catch {
return Response.json({ error: 'Body must be JSON' }, { status: 400 });
}
if (typeof text !== 'string' || text.length < 100) {
return Response.json({ error: 'text must be a string of at least 100 characters' }, { status: 400 });
}
if (text.length > 20_000) {
return Response.json({ error: 'text is too long' }, { status: 413 });
}
try {
const { summary } = await env.AI.run('@cf/facebook/bart-large-cnn', {
input_text: text,
max_length: 150,
});
return Response.json({ summary });
} catch (err) {
console.error('summarization failed', err);
return Response.json({ error: 'Model call failed' }, { status: 502 });
}
},
};
A few decisions here are deliberate. The JSON parse is wrapped separately, so a malformed body returns a 400 instead of a misleading 500. The upper length limit exists because a public endpoint that accepts any text size is an easy way for someone else to spend your inference budget. Summarization models also have a fixed input window, so very long documents need to be chunked and summarized in stages anyway. Model failures map to 502 because the error came from an upstream service, and logging the error keeps it visible in wrangler tail.
If this endpoint is called from a browser on another origin, add CORS headers on both the OPTIONS preflight response and the actual POST response. A common bug is to answer the preflight correctly and then forget the Access-Control-Allow-Origin header on the real response. The browser then blocks the result even though the Worker returned 200.
RAG with Vectorize
Retrieval-augmented generation has two phases: ingest (embed your documents and store the vectors) and query (embed the question, find the nearest chunks, and give them to the LLM as context).
Create the index
npx wrangler vectorize create docs-index --dimensions=768 --metric=cosine
[[vectorize]]
binding = "VECTORIZE"
index_name = "docs-index"
The dimension count must match the embedding model exactly. @cf/baai/bge-base-en-v1.5 produces 768-dimensional vectors; other models produce other sizes. The index’s dimensions and metric cannot be changed after creation.
Ingest inside a Worker
Embeddings can only be computed where the AI binding exists, so ingestion runs as a Worker route (or a scheduled or queue-driven Worker), not as a standalone Node script:
export interface Env {
AI: Ai;
VECTORIZE: VectorizeIndex;
}
async function ingest(env: Env, docs: { id: string; text: string }[]) {
// bge accepts an array of strings, so one call embeds the whole batch
const { data } = await env.AI.run('@cf/baai/bge-base-en-v1.5', {
text: docs.map((d) => d.text),
});
await env.VECTORIZE.upsert(
docs.map((doc, i) => ({
id: doc.id,
values: data[i],
metadata: { text: doc.text },
})),
);
}
Batching matters. Calling the model once per document with Promise.all works, but it makes many separate subrequests, and a Worker has a per-invocation subrequest limit. The array form of text embeds a batch in one call. For a large corpus, push document ids onto a Queue and ingest in batches from a consumer instead of trying to do everything in one HTTP request.
Query
async function answer(env: Env, question: string) {
const { data } = await env.AI.run('@cf/baai/bge-base-en-v1.5', { text: [question] });
const { matches } = await env.VECTORIZE.query(data[0], {
topK: 4,
returnMetadata: 'all',
});
const context = matches
.filter((m) => m.score > 0.5) // tune on real queries; see below
.map((m) => String(m.metadata?.text ?? ''))
.join('\n\n');
const result = await env.AI.run('@cf/meta/llama-3.1-8b-instruct-fp8', {
messages: [
{
role: 'system',
content:
'Answer using only the context below. If the answer is not in the context, say you do not know.\n\n' +
`Context:\n${context}`,
},
{ role: 'user', content: question },
],
max_tokens: 400,
});
return { answer: result.response, sources: matches.map((m) => m.id) };
}
Three details cause most “RAG returns nothing useful” bug reports on Vectorize:
- Metadata is not returned by default. Without
returnMetadata: 'all'(or'indexed'),m.metadatais empty. The code above then builds an empty context, and the model answers from its own training data. There is no error. - Score direction depends on the metric. With cosine, a higher
scoremeans more similar. The0.5threshold is a placeholder: the right cut-off depends on the model and your documents, so choose it by looking at scores for questions you know the answer to. - Filtering needs metadata indexes. To filter a query by a metadata field such as
tenantIdorlang, create a metadata index for that field withwrangler vectorize create-metadata-indexbefore inserting vectors. Vectors inserted earlier are not indexed for that field.
A failure mode I have seen with this setup is a subtle version of the dimension problem: the dimensions match, but the vectors are still incompatible. The bge embedding models accept a pooling option ("mean", the default, or "cls"), and Cloudflare’s own type definitions warn that embeddings from the two modes are not compatible. If ingestion code uses pooling: "cls" and query code uses the default, every insert and query succeeds, and results are simply poor. Put the model id and all embedding options in one shared function used by both ingest and query, and record them next to the index name. When anything about the embedding changes, create a new index and re-embed.
Keep metadata small. Vectorize limits the size of metadata per vector, so long chunks of text are better stored in D1 or R2, with the id in the vector metadata.
Storing conversations in D1
npx wrangler d1 create chat-db
[[d1_databases]]
binding = "DB"
database_name = "chat-db"
database_id = "<id printed by the create command>"
-- migrations/0001_init.sql
CREATE TABLE conversations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id TEXT NOT NULL,
message TEXT NOT NULL,
response TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now'))
);
CREATE INDEX idx_conversations_user ON conversations(user_id, created_at);
npx wrangler d1 migrations apply chat-db --local # for wrangler dev
npx wrangler d1 migrations apply chat-db --remote # for the deployed database
Local and remote D1 are separate databases. Forgetting --remote is the classic reason a deployed Worker fails with no such table, even though everything worked in wrangler dev.
export interface Env {
AI: Ai;
DB: D1Database;
}
export default {
async fetch(request: Request, env: Env, ctx: ExecutionContext): Promise<Response> {
const { userId, message } = await request.json<{ userId: string; message: string }>();
const ai = await env.AI.run('@cf/meta/llama-3.1-8b-instruct-fp8', {
messages: [{ role: 'user', content: message }],
max_tokens: 400,
});
// Do not make the user wait for the log write
ctx.waitUntil(
env.DB.prepare('INSERT INTO conversations (user_id, message, response) VALUES (?, ?, ?)')
.bind(userId, message, ai.response ?? '')
.run(),
);
return Response.json({ response: ai.response });
},
};
ctx.waitUntil lets the insert finish after the response is sent. Always use .bind() with placeholders. Building SQL by string concatenation with user messages is an SQL injection bug like anywhere else. In a real app, userId should come from a verified session or token, not from the request body.
Streaming responses
For chat UIs, streaming gets the first tokens to the user much sooner than waiting for the full completion:
const stream = await env.AI.run('@cf/meta/llama-3.1-8b-instruct-fp8', {
messages,
stream: true,
});
return new Response(stream, {
headers: { 'content-type': 'text/event-stream', 'cache-control': 'no-cache' },
});
With stream: true the result is a ReadableStream of server-sent events, each carrying a small JSON fragment, and the stream ends with a [DONE] event. The client has to parse SSE (for example with EventSource for GET requests, or by reading fetch response bodies line by line). If you also want to save the full answer to D1, tee() the stream and accumulate one branch inside ctx.waitUntil.
Caching identical requests
async function sha256(input: string): Promise<string> {
const buf = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(input));
return [...new Uint8Array(buf)].map((b) => b.toString(16).padStart(2, '0')).join('');
}
const key = `sum:v1:${await sha256(text)}`;
const cached = await env.CACHE.get(key);
if (cached) return Response.json({ summary: cached, cached: true });
const { summary } = await env.AI.run('@cf/facebook/bart-large-cnn', { input_text: text });
ctx.waitUntil(env.CACHE.put(key, summary, { expirationTtl: 86400 }));
Hash the input rather than using the prompt itself as the key: KV keys have a maximum length, and a long prompt used as a key fails. Put a version (v1) in the key so you can invalidate everything by bumping it when you change the model or prompt. Caching makes sense for deterministic tasks such as summarizing the same article or embedding the same text. For open-ended chat, exact-match cache hits are rare.
How the cost adds up
Workers AI bills in Neurons, Cloudflare’s unit for GPU compute, with a daily free allocation and a per-Neuron price beyond it. How many Neurons a request uses depends on the model and on the input and output token counts. Larger models use more per token. It is not tied to the model’s parameter count in any simple way, so calculations like “8B parameters = 8 billion Neurons” are wrong. Check the current per-model rates on Cloudflare’s pricing page, then estimate with your real token counts.
In practice, the main cost drivers are ones you control: max_tokens (without it, a chat model can produce long answers you pay for), how much retrieved context you put in each prompt (four chunks of 500 tokens is 2,000 input tokens on every question), and repeated identical calls that a cache would absorb. Vectorize and D1 have their own usage-based pricing, based on stored and queried vector dimensions and on rows read and written. For most small RAG apps these are small next to inference, but check them when you size an ingest job.