LangSmith

The observability layer for everything LangChain and LangGraph do: every prompt, every tool call, every decision, recorded so you can see exactly where a system went right or wrong.

01

The debugging problem multi-step systems create

Was it the prompt?
Did retrieval pull the wrong documents?
Did the agent choose the wrong tool?
Did the model hallucinate?

A single API call fails loudly: an exception, a stack trace, something to grep for. A multi-step LLM system fails quietly. It just returns a subtly wrong, off-topic, or misleading answer, and every layer along the way, the prompt, the retrieval step, the tool call, the model itself, is a plausible culprit.

Was it the prompt? Did retrieval pull the wrong documents? Did the agent pick the wrong tool? Did the model just hallucinate? Without visibility into each step, answering that means guessing, then re-running the whole pipeline and hoping something stands out.

This gap, observability treated as an afterthought rather than something designed in, is exactly what LangSmith exists to close.

02

Tracing: seeing every step, not just the output

Input
->
Retrieval
->
LLM
->
Tools
->
Output

LangSmith records a full trace of a run: the input prompt, the model's response, which tools were called and what they returned, which documents were retrieved, and any intermediate reasoning, all laid out as a structured, step-by-step timeline.

When a run produces a wrong answer, open its trace and look directly at the failure point instead of re-running the system and staring at the final output. Bad retrieval shows up as the wrong documents in the trace. A bad tool call shows up as a bad tool response. Most of the earlier ambiguity disappears.

It's the same idea as a request trace in any distributed system, applied to a pipeline whose "services" happen to be prompts, retrievers, and model calls instead of microservices.

03

Evaluation: measuring quality instead of guessing

Accuracy
Relevance
Latency
Tokens

LLM output isn't binary, correct or wrong, the way a unit test is. A response can be partially right, off-topic, or subtly misleading, which makes "did that change help?" a genuinely hard question to answer by eye.

LangSmith lets you build evaluation datasets and run a system against them automatically, scoring accuracy, relevance, latency, and token usage. Change a prompt or tweak the retrieval logic, and it's possible to measure whether it actually helped instead of assuming it did.

That same infrastructure supports systematic prompt experiments: running A/B comparisons of prompt variants against the same dataset, so wording changes get evaluated with evidence instead of a gut feeling.

04

Production monitoring, and how the pieces fit

User
->
LangGraph
RouterNodesEdges
->
LangChain
PromptRetrievalModel callTools
->
LangSmith
TraceEvalMonitor

LangSmith logs directly from LangGraph and LangChain as the run happens, not just at the end

Once a system is live, LangSmith keeps tracking it: latency, error rates, token usage, and which tools get invoked under real traffic, the same role a tool like Grafana plays for distributed systems, built specifically for LLM pipelines.

Put together, the three tools split cleanly by concern. LangChain is the vocabulary: the prompts, retrieval, chains, and tools an application is built from. LangGraph is the control flow: the graph of nodes, edges, and state deciding what runs and when. LangSmith is the visibility layer watching the whole thing.

A request comes in, LangGraph decides the execution path, LangChain components do the actual work at each node, prompt construction, retrieval, the model call, the tool execution, and LangSmith records every input, output, and decision along the way. None of the three substitute for good judgment about when this complexity is justified. A single prompt doesn't need any of it. A system that has to be trusted, debugged, and improved over months does, and that's precisely the gap this stack was built to close.

Try it: find the bug in the trace

Same four questions from the start of this page, prompt, retrieval, tool, or model, now with an actual trace to read. Guess before you check.

A user asked: "What's our time-off policy for new hires?" and got the wrong answer. Click the stage of the trace you think caused it.

Waiting for your guess.