LangChain

The plumbing for turning a raw LLM call into a real application: loading data, splitting it into chunks, retrieving the right pieces, and wiring the steps together.

01

From one API call to a pipeline

Prompt
->
LLM
->
Output

a raw API call: one request in, one response back

Talking to an LLM through its raw API is simple: send a prompt, get text back. A document Q&A system needs more than that. It has to find the right document out of potentially thousands, build a prompt that includes the relevant context, call the model, then parse and pass along whatever comes back.

That's no longer a single call, it's a pipeline: retrieve, build a prompt, call the model, parse the output, decide the next step. Every one of those stages needs writing, and doing it from scratch on every project means rewriting the same boilerplate each time.

LangChain exists to give an application that scaffolding as reusable building blocks, so an LLM-powered app, a chatbot, an assistant, a full generative AI system, gets built by composing pieces instead of hand-rolling every stage.

02

Data ingestion: a loader for every source

PDF
Excel / CSV
Website
Wikipedia
arXiv
->
Document Loaderparses each source

Before any of that can happen, the raw material has to get into the system. Document loaders read from wherever the data actually lives: PDFs, Excel and CSV files, websites, Wikipedia, arXiv, an internal wiki, whatever the source is.

This is parsing, and it's the step people underrate most. The more accurately a loader extracts structure and text, the more accurately everything downstream, retrieval, generation, will work. Get this wrong and no amount of clever prompting fixes it later.

LangChain ships a large catalog of loaders because every source has its own quirks. Building a good one for a new source, say scraping a specific site's layout, is itself a real engineering task, and often the highest-leverage one in the whole system.

03

Chunking and the context window

the full document, versus what actually fits the context window

Every LLM has a context window, a hard limit on how much text fits into a single call. A full document, or a whole knowledge base, will not fit.

Text splitters solve this by breaking documents into smaller chunks before anything is stored or retrieved. Several strategies exist for where to cut, by character count, by sentence, by semantic boundary, and the choice matters: cut in the wrong place and the exact sentence that answered the question gets severed in two.

Those chunks are then converted into vectors, an embedding capturing meaning rather than exact words, and stored in a vector database for retrieval later. That conversion step is what makes semantic search possible in the first place.

04

Retrieval: the right context, not just any context

Query"how much PTO do I get"
->
Vector DBsemantic search
->
Prompt + LLMcontext injected
->
Answergrounded in retrieved text

When a user asks a question, the system does not search the raw documents, it searches the vector database, using the query's own embedding to find chunks that are semantically close, not just textually similar. That is what makes it find "vacation policy" when the user typed "how much PTO do I get."

Those retrieved chunks become the context injected into the prompt sent to the LLM. Get retrieval wrong, the wrong chunks, missing chunks, and the model will confidently answer from the wrong material. This is usually where a shaky system actually breaks, not in the model itself.

This whole retrieve-then-generate pattern has a name: RAG. The model does not need to know everything, it just needs the right information handed to it at the right time.

05

Chains, tools, and where it strains

Prompt
->
LLM
->
Context
->
Output

Once retrieval works, LangChain wires the remaining steps into a chain: a prompt template, the LLM, the retrieved context, each one feeding the next, executed strictly in order (in current LangChain this composition happens through LCEL, chaining Runnables together with a pipe operator). A chain can summarize a document, extract key points, then generate a final answer. It can even branch to a different prompt depending on the input. What it cannot do is loop back: once a chain reaches its last step, there's no path back to an earlier one within that same run.

Tools extend what any single step can do. By default an LLM can only generate text, it cannot run a calculation, check a live price, or query a database. A tool is a function the model can ask the framework to invoke on its own behalf, and the result comes back into the model's context so it can keep reasoning with real, current data.

The honest tradeoff: for a single LLM call, all of this is pure overhead, hit the API directly instead. And its layers mean that when something breaks, debugging often means working through the framework before reaching your own logic. It earns its keep once there is an actual multi-step pipeline to manage. Its sharpest limit is that loop: the moment a workflow needs to retry, re-plan, or send execution back to an earlier step, a chain has no way to express that, which is exactly the gap LangGraph was built to close.

Try it: watch retrieval happen

Press run and follow one question through the whole pipeline: embedded, searched, matched against stored chunks, and answered from what came back.

Query
->
Embed
->
Search
->
Chunks
->
Prompt
->
Answer

Press run to send a query.