Tag: LLM
5 posts
-
Running LLMs Locally with Ollama: Quantization Tags, num_ctx, GPU Offload and the API
Run LLMs locally with Ollama: what quantization tags mean, why num_ctx silently truncates prompts, GPU/CPU offload, the OpenAI-compatible API, and safe network exposure.
-
LangChain 1.x in Practice: LCEL Pipes, create_agent, RAG and When to Skip It
Current LangChain in Python: the package split, LCEL Runnables, memory that actually works, create_agent replacing AgentExecutor, a RAG chain, and when to skip it.
-
Building on the Claude Messages API: max_tokens, Streaming, Tool Use Loops, Caching and Retries
Claude Messages API in production: required max_tokens, content blocks, streaming, the tool use loop, prompt caching, 429/529 retries and stop_reason values.
-
Running AI Models at the Edge with Workers AI: Summarization API, Vectorize RAG and D1
Workers AI in practice: bindings, a summarization endpoint, RAG with Vectorize, logging to D1, streaming, and the pitfalls around dimensions, metadata and cost.
-
OpenAI API in Code: Chat Completions, Responses, Tool Calls, Structured Outputs and Retries
The OpenAI API with the current Python and Node SDKs: Chat Completions vs Responses, tools and tool_choice, Structured Outputs, streaming, and 429 retries.