LangChain 1.x in Practice: LCEL Pipes, create_agent, RAG and When to Skip It

Key takeaways

LangChain's API has changed several times, and most tutorials online still show code that no longer imports. This guide uses the 1.x layout: langchain-core Runnables and LCEL pipes for fixed pipelines, create_agent for tool calling, checkpointers for memory, plus an honest look at when a plain SDK call is the better choice.

Why most LangChain tutorials break on first run

LangChain moves faster than almost any other Python library you’re likely to depend on. The shape of the “hello world” has changed several times: LLMChain and ConversationChain, then LCEL pipes, then AgentExecutor, and now LangGraph-based agents. A large share of the code on blogs and Stack Overflow no longer imports. On a fresh install of LangChain 1.x, the classic snippets fail like this:

>>> from langchain.chains import LLMChain
ModuleNotFoundError: No module named 'langchain.chains'

>>> from langchain.agents import AgentExecutor
ImportError: cannot import name 'AgentExecutor' from 'langchain.agents'

>>> from langchain.memory import ConversationBufferMemory
ModuleNotFoundError: No module named 'langchain.memory'

I checked each of these against langchain 1.4 / langchain-core 1.6. initialize_agent and create_tool_calling_agent are also gone from langchain.agents. The legacy APIs live on in a separate langchain-classic package for code you can’t migrate yet. New code should use the patterns below.

The package split

LangChain is no longer one package. Knowing what lives where saves a lot of confusion:

PackageWhat’s in it
langchain-coreThe Runnable interface, prompts, messages, output parsers, @tool, base vector store/retriever classes
langchain-openai, langchain-anthropic, langchain-ollama, …One package per provider: chat models and embeddings
langchainHigher-level building blocks: create_agent, init_chat_model, agent middleware
langchain-text-splittersDocument chunking
langchain-communityLong-tail third-party integrations, many of which have moved to dedicated packages (e.g. langchain-chroma)
langgraphThe graph runtime that create_agent is built on: state, checkpointers, human-in-the-loop
pip install langchain langchain-openai langchain-text-splitters
export OPENAI_API_KEY="sk-..."

Pin versions in requirements.txt. Deprecations in LangChain arrive in minor releases, and a pipeline that emits warnings today may stop importing a few versions later.

LCEL: prompt | model | parser

The core abstraction is the Runnable. Prompts, models, parsers, retrievers and plain functions all expose the same methods: invoke, batch, stream, and their async versions ainvoke, abatch, astream. The | operator composes them into a RunnableSequence. This is the LangChain Expression Language (LCEL).

from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

prompt = ChatPromptTemplate.from_messages([
    ("system", "You explain technical concepts in two sentences."),
    ("human", "{question}"),
])

chain = prompt | llm | StrOutputParser()

print(chain.invoke({"question": "What is a vector database?"}))

# The same chain, for free:
for token in chain.stream({"question": "What is RAG?"}):
    print(token, end="", flush=True)

answers = chain.batch([{"question": "What is BM25?"}, {"question": "What is HNSW?"}])

What the pipe buys you is that every step has the same interface. So stream() passes tokens through the parser, batch() runs requests concurrently, and retries or fallbacks can wrap any step:

from langchain_anthropic import ChatAnthropic

robust_llm = llm.with_retry(stop_after_attempt=3).with_fallbacks(
    [ChatAnthropic(model="claude-sonnet-4-5")]  # use whatever model name your account supports
)
chain = prompt | robust_llm | StrOutputParser()

The cost is that errors come from inside the framework. A misnamed input key, for example, doesn’t fail where you wrote it. It fails in the prompt:

KeyError: "Input to ChatPromptTemplate is missing variables {'question'}.
Expected: ['question'] Received: ['q']"

That message is at least clear. Errors from deeper inside a long sequence often aren’t, which is why tracing (below) matters more with LangChain than with raw SDK calls.

Plain functions become steps via RunnableLambda, and dict literals in a pipe become RunnableParallel, which runs its branches concurrently and collects the results into a dict. The RAG chain below uses exactly that.

Structured output

When you need typed data instead of prose, don’t parse JSON out of free text. Use the model’s native structured output support:

from pydantic import BaseModel, Field

class Product(BaseModel):
    name: str
    price: float = Field(description="Price in USD")
    in_stock: bool

extractor = llm.with_structured_output(Product)
item = extractor.invoke("Widget Pro, $29.99, currently available")
print(item.name, item.price, item.in_stock)

Under the hood this uses the provider’s tool-calling or JSON-schema mode, so the reliability depends on the model. Small local models in particular produce schema violations more often. Catch validation errors and decide whether to retry or fail.

Memory: the bug in half the examples

Conversation memory is where copy-pasted examples most often silently fail. The typical RunnableWithMessageHistory example stores history per session, but the prompt has no slot for it:

prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a helpful assistant."),
    ("human", "{question}"),   # no history placeholder!
])

I instrumented the model to log which messages it actually received. On the second turn it got only ['system', 'human'], so the “My name is Alice” turn was stored but never sent. Adding MessagesPlaceholder("history") (matching history_messages_key) changed that to ['system', 'human', 'ai', 'human']. No error, no warning, just a bot that can’t remember anything.

In current releases, RunnableWithMessageHistory and InMemoryChatMessageHistory emit deprecation warnings pointing to LangGraph persistence. The current approach is a checkpointer keyed by thread_id:

from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver

assistant = create_agent(
    llm,
    tools=[],
    system_prompt="You are a helpful assistant.",
    checkpointer=InMemorySaver(),
)

config = {"configurable": {"thread_id": "user-123"}}
assistant.invoke({"messages": [{"role": "user", "content": "My name is Alice."}]}, config)
reply = assistant.invoke({"messages": [{"role": "user", "content": "What's my name?"}]}, config)
print(reply["messages"][-1].content)

Each thread_id keeps its own message list. InMemorySaver loses everything on restart and isn’t shared across workers. In production, use a database-backed checkpointer (LangGraph provides Postgres and SQLite ones as separate packages).

Unbounded history is the next problem. Every turn resends the entire conversation, so cost and latency grow with conversation length until you hit the context limit. Decide early whether you trim to the last N messages or summarize older ones.

Agents with create_agent

An agent is a loop: the model either answers or requests tool calls, the runtime executes the tools, appends the results, and calls the model again. In LangChain 1.x that loop is create_agent, built on LangGraph:

from langchain.agents import create_agent
from langchain_core.tools import tool

@tool
def get_order_status(order_id: str) -> str:
    """Look up the shipping status of an order by its ID, e.g. 'A-1042'."""
    # call your real API here
    return f"Order {order_id}: shipped, arriving Thursday"

agent = create_agent(
    llm,
    tools=[get_order_status],
    system_prompt="You are a support assistant. Use tools for order questions; never guess order data.",
)

result = agent.invoke({"messages": [{"role": "user", "content": "Where is order A-1042?"}]})
for m in result["messages"]:
    print(m.type, m.content)

With a scripted model standing in for the LLM, the message trace looks like this: human → ai (empty content, with a tool call) → tool (the function’s return value) → ai (the final answer). You get the whole trace back, which is useful for debugging and for showing “looking up your order…” states in a UI.

Things that matter more than the framework:

  • The docstring is the tool’s spec. The model decides when to call the tool and how to fill in its arguments from the name, docstring and type hints. Vague docstrings produce wrong calls.
  • Tools are an attack surface. Tool arguments come from model output, which can be steered by user input or by retrieved documents (prompt injection). Validate arguments and authorize them against the current user, not against whatever ID the model produced.
  • Bound the loop. A confused model can call the same tool repeatedly. LangGraph enforces a recursion limit (you’ll see GraphRecursionError when it’s hit), and you can set it per call through the recursion_limit config key. Set it deliberately rather than relying on the default.

For a fixed sequence of steps, don’t use an agent. If you know the flow is always “retrieve, then answer,” a chain is cheaper, faster and deterministic. An agent is for when the model genuinely needs to choose between actions.

A RAG chain you can reason about

Retrieval-augmented generation means fetching relevant chunks of your documents and putting them in the prompt. Here’s the whole pipeline with an in-memory store (swap in Chroma, pgvector, etc. for persistence):

from langchain_openai import OpenAIEmbeddings
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_core.documents import Document
from langchain_core.runnables import RunnablePassthrough
from langchain_text_splitters import RecursiveCharacterTextSplitter

docs = [Document(page_content=open("refund-policy.md").read(), metadata={"source": "refund-policy.md"})]

splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=100)
chunks = splitter.split_documents(docs)          # metadata is copied onto every chunk

store = InMemoryVectorStore.from_documents(chunks, OpenAIEmbeddings(model="text-embedding-3-small"))
retriever = store.as_retriever(search_kwargs={"k": 4})

def format_docs(docs):
    return "\n\n".join(f"[{d.metadata['source']}]\n{d.page_content}" for d in docs)

rag_prompt = ChatPromptTemplate.from_template(
    "Answer using only the context below. If the answer isn't there, say you don't know.\n"
    "Cite the [source] you used.\n\n{context}\n\nQuestion: {question}"
)

rag_chain = (
    {"context": retriever | format_docs, "question": RunnablePassthrough()}
    | rag_prompt
    | llm
    | StrOutputParser()
)

print(rag_chain.invoke("How long do customers have to request a refund?"))

The dict at the start is a RunnableParallel: the question string goes both to the retriever (then format_docs) and straight through (RunnablePassthrough). A single-variable prompt also accepts a bare string as input, which is why invoke("...") works here.

Most RAG quality problems are retrieval problems, not LangChain problems:

  • Chunk size is a trade-off. Small chunks match precisely but lose surrounding context. Large chunks keep context but dilute the embedding and use up the prompt. Look at what the retriever actually returns for real questions before tuning the prompt.
  • Embedding models aren’t interchangeable. Vectors from one embedding model are meaningless to another. If you change models, re-embed the whole corpus, or queries will return garbage with no error.
  • Pure vector search misses exact terms such as error codes, SKUs and function names. Hybrid search (vector plus keyword/BM25) or metadata filters often help more than a better LLM.
  • “Only use the context” is a request, not a guarantee. Keep citations, so users and you can check answers against their sources.

For the storage side, see Local Vector Search with ChromaDB.

Swapping providers, including local models

Provider independence is the strongest practical argument for LangChain. The chain above doesn’t care which chat model it gets:

from langchain.chat_models import init_chat_model

llm = init_chat_model("openai:gpt-4o-mini", temperature=0)
# llm = init_chat_model("anthropic:claude-sonnet-4-5")    # needs langchain-anthropic
# llm = init_chat_model("ollama:llama3.1")                # needs langchain-ollama and a running Ollama

“Swappable” has limits, though. Prompts tuned for one model often behave differently on another. Small local models are weaker at tool calling and structured output. And provider-specific features (prompt caching, extended thinking, some multimodal inputs) are exposed unevenly. Re-run your evaluation set after switching, don’t just change one string. For running models locally, see Running LLMs Locally with Ollama.

Observability with LangSmith

Because a single invoke can fan out into retrievals, several model calls and tool executions, tracing is close to mandatory once you go past one step:

export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="..."
export LANGSMITH_PROJECT="support-bot"   # optional: groups traces

Every Runnable and agent step then shows up as a nested trace with inputs, outputs, latency and token usage. Be aware of what this means for data handling: full prompts and retrieved documents are sent to LangSmith. If your documents are sensitive, check your organization’s policy first or self-host.

When LangChain adds more than it saves

I’ve seen the same arc on several projects. A team adopts LangChain for a feature that’s really “one prompt, one call.” Six months later, a minor upgrade renames imports and floods the logs with deprecation warnings. A bug surfaces from inside a RunnableSequence stack trace, and someone spends an afternoon reading framework source to find that the prompt was missing a variable. The provider SDK version of that feature would have been fifteen lines with nothing between you and the HTTP request. For single-call features, I now start with the SDK (see the Claude API guide) and reach for LangChain only when composition actually shows up.

It earns its place when you have several of these at once:

  • retrieval + prompt + parsing + fallbacks composed into one streamed pipeline,
  • a real need to switch or mix providers (cloud and local),
  • multi-step tool-calling agents that need persistence, human approval steps or tracing,
  • a team that benefits from shared, standard building blocks.

If you adopt it, keep the LangChain layer thin: business logic in plain functions, LangChain as glue. That makes the next API migration a one-file change instead of a rewrite.