~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / langgraph-mem0-langchain-mcp-what-you-actually-need $
Tutorials

LangGraph, Mem0, LangChain, MCP: What You Actually Need

LS Lucas Souza · · 15 min read
LangGraph, Mem0, LangChain, MCP: What You Actually Need

You have 6 documentation tabs open and you still haven't written a line of code. LangGraph in one, Mem0 in another, LangChain, MCP, a vector store comparison and a 40-minute YouTube tutorial. The problem isn't picking the wrong tool. It's assuming you need one.

The orchestration and memory ecosystem for agents matured too fast and nobody stopped to write the honest map. Every project has a landing page that says "build production-ready agents" and none of them tells you the exact point where it starts pulling its own weight. This post is that map.

You'll walk away with four marketing-free definitions, the cutoff criterion for each tool (the moment it stops being overhead and becomes infrastructure) and a 5-question decision tree to run before installing anything. Including the case — a common one — where the right answer is "none of them".

TL;DR

  • What it is: a decision map of the agent orchestration (LangGraph, LangChain) and memory (Mem0) ecosystem, plus the real state of MCP.
  • Stack/Models: agnostic. The examples are Python, but the criteria apply just the same to PHP, TypeScript or Go.
  • Cost/Access: LangGraph, LangChain and Mem0 OSS are open source and free. Mem0 Platform starts at $19/month. MCP is an open protocol.
  • Rule of thumb: write the raw loop first. Adopt a framework when you hit a specific pain that has a name, not on principle.

What each one does, in one sentence, no marketing

Before deciding anything, one thing has to be clear: these four things don't compete with each other. Three are libraries at different layers and one is a protocol. Comparing LangGraph to MCP is comparing Laravel to HTTP.

LangChain is the convenience layer. A unified interface for the model, tools, structured output and a ready-made agent loop. In version 1.0, released in October 2025, the library threw out initialize_agent and AgentExecutor and collapsed everything into a single entry point, create_agent, with middleware hooks (before_model, wrap_model_call, wrap_tool_call, after_model) so you can tweak the loop without rewriting it. It's less of a framework and more of a lean harness than the version that traumatized you in 2023.

LangGraph is the runtime. You declare nodes, edges and state; it executes the graph, saves a checkpoint at every node transition and knows how to resume from where it stopped. It's the piece that turns "my script died halfway through" into "my agent paused and picked back up".

Mem0 is a long-term memory layer. It reads the conversation, extracts facts about the user, decides what's new, what conflicts with what it already knew and what to discard — and hands back only the relevant pieces on the next call, instead of you resending the entire history.

MCP isn't a library at all. It's the protocol that standardizes how an AI client discovers and calls tools exposed by a server. It doesn't orchestrate, doesn't store memory, doesn't decide anything. It just replaces custom integration with a common interface.

Notice the pattern: none of the four solves the hard problem. The hard problem is deciding what goes into the context, when the model calls a tool, how you measure whether the answer was any good and how much that costs per request. No tool decides that for you — and that exact set of decisions is what we break down live at the AI Engineering Lab 3ª Edição, a two-day immersion on September 19 and 20 building agent architecture with tool calling, routing, memory, grounding, tracing, evals and cost all on the table. The rest of this post is the map of when each library helps with that design and when it just gets in the way.

Orchestration: when LangGraph pulls its own weight

The most honest recommendation about agent frameworks came from Anthropic itself, in Building Effective Agents: "Developers start by using LLM APIs directly: many patterns can be implemented in a few lines of code." The warning that comes with it is what matters — a framework makes getting started easier and obscures what's running underneath, which makes debugging harder.

A tool-calling agent is literally this:

messages = [{"role": "user", "content": pergunta}]

while True:
    resp = client.messages.create(
        model="claude-sonnet-5",
        tools=TOOLS,
        messages=messages,
    )
    messages.append({"role": "assistant", "content": resp.content})

    if resp.stop_reason != "tool_use":
        break

    resultados = [executar(bloco) for bloco in resp.content if bloco.type == "tool_use"]
    messages.append({"role": "user", "content": resultados})

That's an agent. It's not a prototype, it's not a simplified version for teaching purposes. It's the loop running in production in a lot more places than the landing pages suggest. If your case is "call the model, it picks a tool, run it, respond", installing LangGraph on top of that means adding a graph, a state schema and a checkpointer to keep running the same while.

So when does LangGraph pull its own weight? When you hit one of these four pains, each with a name and a symptom:

  1. Durable execution. The process crashed in the middle of a 12-step sequence and you need to resume from step 9, not from scratch. LangGraph checkpoints at every node transition into a pluggable persistence layer — in-memory for dev, SQLite for something light, Postgres in production. Reimplementing that by hand is doable. It's also real work.
  2. Real human-in-the-loop. The flow needs to stop, wait for a human approval that might show up 3 days later, and continue with its state intact. Without a persisted checkpoint, that turns into a queue hack.
  3. Conditional routing with cycles. It's not if/else; it's a graph where the validation node can send things back to the generation node N times until it passes, with a recursion limit.
  4. Parallel fan-out with state merge. Three nodes running at the same time and writing to the same state object without corrupting anything.

If you can't point to which of these four pains you're feeling right now, the answer is while. And it's not wasted work: the raw loop is where you isolate the assets that are actually yours — the tools, the prompts, the decision logic. A framework plugs into that core later.

The other side is worth putting on record. LangGraph isn't Twitter hype: it has more than 33,000 GitHub stars and tens of millions of monthly downloads on PyPI, with Uber, LinkedIn and Klarna running it in production. The question was never whether it works. It's whether your problem has the shape it solves.

Memory: Mem0 vs. a table in Postgres

The confusion is bigger here, because "memory" has become an umbrella word for three different things.

Short-term memory is the message array of the current conversation. That's a JSON column. No tool required.

Semantic memory is "retrieve the chunks similar to the question". That's vector search, and the Postgres you already have handles it — I covered the full setup in pgvector in Postgres: where to store your agent's memory. I won't repeat the argument here: in short, at the scale where 90% of products live, a table with a vector column and an HNSW index is the right answer.

Memory of facts about the user is where Mem0 comes in and where SELECT alone doesn't cut it. The problem isn't storing, it's keeping it coherent. The user said in March that they're vegetarian and in August that they're eating meat again. You have two contradictory facts in the database. Which one goes into the prompt?

That's the real work: extracting facts from conversation, detecting conflict, resolving conflict, deciding relevance at retrieval time. It's the part that's a pain to write and an even bigger pain to maintain.

The numbers Mem0 publishes on its research page are good: 92.5 on LoCoMo, 94.4 on LongMemEval, and — the number that matters most at the end of the month — around 1,800 tokens per conversation versus 26,000 for sending the whole context. That's a ~90% reduction in cost per call, with p50 latency under 1.1 s. These are the vendor's own numbers, so treat them as a ceiling, not a guarantee. But the order of magnitude makes sense: sending 5 relevant facts is obviously cheaper than sending 40 conversation turns.

The detail almost nobody reads before deciding: Mem0 open source and Mem0 Platform are not the same product. According to the official comparison docs, the OSS version (Apache 2.0) has no graph memory, memory decay, temporal reasoning, batch operations, webhooks or app_id for multi-tenant separation. Passing a decay parameter in OSS returns an error. Self-hosting is free in license terms and costs you three containers — FastAPI, Postgres with pgvector and Neo4j — plus your LLM cost, because fact extraction is a model call on every write. The Platform starts at $19/month and graph memory only shows up in the $249 plan.

So, the cutoff criteria:

  • A table in Postgres if your agent's memory is "what did this user ask before" and "which documents are similar to this".
  • Mem0 (or equivalent) if you need the agent to know facts about the person that span sessions, contradict each other over time and need to be reconciled without intervention.
  • Nothing if your session is stateless. An FAQ chatbot doesn't need to remember anything, and poorly calibrated memory poisons answers with stale context.
▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

MCP: what changed and what's still hype

MCP is the item on the list that changed the most in the last few months, and it changed in a direction almost nobody talked about.

The 2026-07-28 spec is the biggest revision since the protocol launched, and the summary is: MCP went stateless. The initialize/initialized handshake and the Mcp-Session-Id header are gone. Each request carries its own protocol version, identity and capabilities in _meta. In practice, that means any request can land on any server instance behind a round-robin load balancer, with no shared state. MCP stopped requiring special infrastructure and now runs on plain HTTP.

Along with that came changes that matter to anyone operating this stuff:

  • Header-based routing. Mcp-Method and Mcp-Name are now required, so a gateway can route and meter without parsing JSON.
  • Cacheable results. tools/list, prompts/list, resources/list and resources/read now return ttlMs and cacheScope.
  • Tasks left the core and became an official extension (io.modelcontextprotocol/tasks), with tasks/get via polling.
  • Deprecations with a deadline. Dynamic Client Registration gives way to CIMD; Roots, Sampling, Logging and the legacy HTTP+SSE transport have at least 12 months of runway before they disappear.

What's still hype: the idea that plugging in more MCP servers makes the agent better. It makes it worse. Tool definitions take up context, and context is a scarce resource. Anthropic itself showed the size of the problem in Code execution with MCP: in a worked example, swapping direct tool calls for code that loads only the needed definitions on demand took consumption from 150,000 to 2,000 tokens — a 98.7% reduction.

Read that again. The bottleneck wasn't the model. It was the tool catalog you shoved into the window before it ever read the question.

The practical conclusion is boring and correct: MCP is excellent for integration — no more writing a custom adapter for every system, reusing a server that already exists, standardizing auth. It's terrible as a capability strategy. Twenty connected servers isn't a powerful agent, it's an agent with a clogged context.

The decision tree: 5 questions before installing anything

Run them in order. Stop at the first "no".

1. Does the flow have a cycle, conditional branching or parallelism? If it's a linear sequence of steps, you don't need a graph. You need a function. No → no LangGraph.

2. Is a crash in the middle of execution expensive? Expensive = losing minutes of work, money spent on tokens or an action partially applied to an external system. Yes → the LangGraph checkpointer pulls its own weight. No, it's a 3-second request → a retry handles it.

3. Does someone need to approve something along the way? Human-in-the-loop with a long wait requires persisted state and resumption. That's LangGraph's specialty and it's tedious to do by hand. No → stick with the loop.

4. Does the agent need to remember facts that span sessions and contradict each other over time? No, I just need similar chunks → pgvector. No, it's stateless → nothing. Yes → now you want a dedicated memory layer, with the OSS vs. Platform caveat.

5. Are you integrating with a system that already has an MCP server ready? Yes → use MCP, it's real code savings. No, it's your own API → a Python function with a JSON schema is simpler, faster and cheaper on context than standing up an MCP server to talk to yourself.

Five "no"s is a legitimate result and more common than it seems. It means: model SDK, a while, your functions, your Postgres. That stack holds up in production.

And there's question zero, which comes before the five: do you already have tracing? If you can't reconstruct why the agent decided what it decided, switching frameworks won't fix anything — it'll just move your blind spot somewhere else. That part has its own post in AI agent observability.

Limitations and things to watch out for

A framework isn't neutral on context cost. An abstraction that builds the prompt for you builds a prompt you haven't read. Before blaming the model for a bad answer, log the final payload that went out to the API. I've seen a system prompt doubled in size by middleware that was "helping".

Managed memory is a leak surface. Sending customer conversations to a third-party memory service is personal data processing. Under LGPD, that means a legal basis, a data processor agreement, a retention policy and a deletion path that actually works. If that conversation hasn't happened on your team yet, self-hosting solves the legal problem and hands you back the infra problem.

MCP widens the attack surface. A third-party MCP server is code that describes tools to your model — and a tool description is text that goes into the context. Prompt injection via tool description is a real vector and has already earned a dedicated security guide from the NSA. Treat an external MCP server with the same care as a dependency with write permission.

A vendor benchmark measures the vendor's case. Mem0's numbers are Mem0's. They're probably right for the scenario they tested. Your scenario is a different one. Measure on your own data before signing up for any plan.

Quick FAQ

Are LangChain and LangGraph the same thing? No. LangGraph is the graph runtime with state and checkpoints; LangChain is the convenience layer (model, tools, create_agent, middleware) that runs on top of it. You can use LangGraph on its own. You can use LangChain without ever drawing a graph.

Do I need LangChain to use MCP? No. MCP is a protocol with official SDKs in TypeScript, Python, Go and C#. You speak MCP straight from your code, with no agent framework in between.

Is it worth migrating an agent that already runs on a plain while to LangGraph? LangGraph is worth it once you've felt one of the four pains: durable execution, human-in-the-loop with a long wait, conditional cycles or parallel fan-out. Migrating "to look more professional" is refactoring without a hypothesis — you trade code you understand for an abstraction you don't understand yet.

What's the best alternative to LangChain? Depends on what bothers you about it. If it's the weight of the abstraction, the alternative is the model SDK directly — Anthropic, OpenAI or Gemini — with your own functions. If it's the ecosystem, there are leaner, typed options like Pydantic AI, and in the PHP world Prism covers the same ground. There's no single replacement because LangChain doesn't do just one thing.

Is Mem0 OSS enough or do I need the Platform? It's enough if you want fact extraction and reconciliation with your own storage. If you need graph memory, memory decay or temporal reasoning, those features don't exist in OSS and the parameter returns an error — it's either the Platform or your own implementation.

The takeaway

The ecosystem isn't confusing. It's layered, and the layers have become clear: protocol (MCP), runtime (LangGraph), convenience (LangChain), memory (Mem0 and competitors). What's confusing is the assumption, which nobody wrote down but everybody carries around, that you need a piece at every layer to be taken seriously.

You don't. You need a loop you understand, tools you've tested, context you control and a number that tells you whether it's working.

Write the while first. Run it in production. When it hurts, the name of the pain will tell you exactly which piece to install — and you'll install one, not six. That's the difference between choosing an architecture and collecting dependencies.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing