Tag: AI
11 posts
-
Running LLMs Locally with Ollama: Quantization Tags, num_ctx, GPU Offload and the API
Run LLMs locally with Ollama: what quantization tags mean, why num_ctx silently truncates prompts, GPU/CPU offload, the OpenAI-compatible API, and safe network exposure.
-
LangChain 1.x in Practice: LCEL Pipes, create_agent, RAG and When to Skip It
Current LangChain in Python: the package split, LCEL Runnables, memory that actually works, create_agent replacing AgentExecutor, a RAG chain, and when to skip it.
-
Cursor vs Copilot vs Claude Code vs Windsurf vs Aider: How the AI Coding Tools Differ
A comparison of top AI coding assistants in 2026, covering code completion quality, context handling, and pricing.
-
Building MCP Servers: Model Context Protocol Architecture, Python and TypeScript Servers, Claude Desktop
MCP (Model Context Protocol) is Anthropic's open standard for connecting AI models to external tools. How it is structured, building servers in Python and TypeScript, and connecting them to Claude Desktop.
-
Building on the Claude Messages API: max_tokens, Streaming, Tool Use Loops, Caching and Retries
Claude Messages API in production: required max_tokens, content blocks, streaming, the tool use loop, prompt caching, 429/529 retries and stop_reason values.
-
Working in Cursor: Tab Completion, Chat, Agent Edits and Project Rules Compared with VS Code + Copilot
How Cursor's Tab completion, Chat, multi-file agent edits and project rules behave, where they go wrong, and how Cursor compares with VS Code plus Copilot.
-
Running AI Models at the Edge with Workers AI: Summarization API, Vectorize RAG and D1
Workers AI in practice: bindings, a summarization endpoint, RAG with Vectorize, logging to D1, streaming, and the pitfalls around dimensions, metadata and cost.
-
OpenAI API in Code: Chat Completions, Responses, Tool Calls, Structured Outputs and Retries
The OpenAI API with the current Python and Node SDKs: Chat Completions vs Responses, tools and tool_choice, Structured Outputs, streaming, and 429 retries.
-
Local Vector Search with ChromaDB: Embeddings, Queries, LangChain and a RAG Chatbot
Local vector search with ChromaDB 1.x: how embeddings and distances work, metadata filters, persistence traps, LangChain integration, and a minimal RAG chain.
-
Prompt Engineering for Developers: Few-Shot, Chain-of-Thought and Constraint Prompts for Code
Get usable code out of ChatGPT, Claude and Cursor with few-shot, chain-of-thought, role and constraint prompts, plus fixes for context loss and bloated output.
-
Python Meets C++: High-Performance Engines with pybind11
Bind C++ to Python with pybind11: minimal modules, CMake/setuptools builds, NumPy buffers, GIL release, wheels, and production patterns.