Free forever, no credit card.Get Started for Free →
← All posts
October 8, 2026 · 6 min read

Pipecat Remembers the Pipeline. Your Callers Start From Zero Every Call.

Pipecat Remembers the Pipeline. Your Callers Start From Zero Every Call. A voice AI agency ships an appointment-setter for a dental clinic. The agent calls a patient, books a cleaning, hangs up. Two days later it calls the same patient to confirm, and opens with the full intake script again: name, date of birth, insurance, reason for visit. The patient sighs and says, "I gave you all of this on Tuesday." Nothing broke. The pipeline ran exactly as built. Pipecat did everything it was designed t

Pipecat Remembers the Pipeline. Your Callers Start From Zero Every Call.

A voice AI agency ships an appointment-setter for a dental clinic. The agent calls a patient, books a cleaning, hangs up. Two days later it calls the same patient to confirm, and opens with the full intake script again: name, date of birth, insurance, reason for visit. The patient sighs and says, "I gave you all of this on Tuesday."

Nothing broke. The pipeline ran exactly as built. Pipecat did everything it was designed to do, and none of what it was designed to do includes remembering Tuesday.

What Pipecat is actually good at

Pipecat is an open-source Python framework for real-time voice (and multimodal) AI pipelines. It handles the hard real-time parts: streaming audio in, speech-to-text, feeding turns to an LLM, streaming text-to-speech back out, and managing the transports that carry the call. It works with transports from Daily, Deepgram, OpenAI, and Cartesia, and with whichever LLM provider you point it at.

Think of it as the plumbing of a live conversation: the pipeline assembles per call and tears down when the call ends.

What persists between calls: nothing

Pipecat is stateless by design. Each call spins up its own pipeline and its own LLM context. When the session ends, nothing is persisted: no conversation history, no facts about the caller, no record of what was decided. The next call gets a fresh system prompt and a fresh context, as if the caller has never existed.

Your code persists, of course. The pipeline definition, the system prompt, the tool wiring, all of that lives in your repo. But everything that happened during the call dies with the pipeline. Pipecat remembers how the pipeline is built. It remembers nothing about the people who called through it.

The session is not the caller

Inside a call, Pipecat does maintain short-term context. Context aggregators collect the turns of the current conversation so the LLM can respond coherently mid-call. That is genuine in-call memory, and it works well for what it covers: the last few minutes of dialogue.

The trap is assuming that coverage extends. It does not extend past hang-up, and it does not extend to anything the caller said last week. An operator who confuses in-call context with caller memory discovers the gap the way the dental clinic did: the agent is fluent during the call and amnesiac between calls.

Why stateless is a design choice, not a bug

Statelessness is what lets a voice pipeline scale. No shared mutable state between calls means any worker can handle any call, retries are safe, and a crashed session leaves nothing half-written behind. Persistence was deliberately left out of the framework's job description, which leaves the operator with an integration task: deciding what gets remembered, whose memory it is, when it gets read, and when it gets written.

The four decisions every Pipecat memory layer has to make

1. Identity. Memory needs an owner. The standard approach is one memory bank per caller: the phone number, or better, an authenticated account or customer ID when one person calls from multiple numbers. Get normalization wrong and the same caller splits into two strangers, or two callers merge into one. The identity key is the single most consequential design decision in the whole layer.

2. Recall before the LLM. Memory has to be fetched and injected into the context before the LLM responds. That fetch sits on the critical path of a real-time conversation, so it needs a hard latency cap with fail-open behavior: if recall is slow, the call continues without the memory rather than making the caller wait.

3. Retain after each turn. What gets saved, and when? The workable pattern is retaining the substance of completed turn pairs: requests, decisions, stated preferences, commitments. Not the full transcript, the distilled facts. Retain runs asynchronously and never blocks the response path, because nothing the caller needs depends on the write finishing.

4. Scope. What the agent remembers about a caller should be the caller's business with you, not a replay of every word. Preferences, history, open issues, prior commitments. The memory that earns its keep is the kind that changes what the agent does next time.

Your options, honestly

You can build the layer yourself: a Postgres table keyed by caller, a recall function that runs before the LLM call, a retain function that writes after each turn. It works, and it is the option with the fewest surprises, at the cost of being yours to maintain, scale, and secure.

Or you can use a purpose-built memory service. Mem0 and Zep are general memory APIs you wire into the pipeline yourself. Hindsight is built specifically for Pipecat: a single frame processor between the user aggregator and the LLM that recalls before inference and retains after each turn. Chanl pairs memory with a knowledge base, tool management over MCP, and prompt versioning in one SDK, with wrappers for Pipecat and LiveKit. InfoLang offers a semantic-memory processor you drop into the same slot.

All of them solve the same core problem, because Pipecat itself chose not to. The choice is about how much of the surrounding stack you want one vendor to own.

One memory for the whole operation

There is a wrinkle the dedicated voice-memory tools do not cover. Your voice agent is rarely your only agent. The same business often runs scheduled agents too: a nightly follow-up agent, an n8n workflow that syncs bookings, a research agent that preps call lists. If the voice agent learns that a patient prefers morning appointments and that fact lives only in the voice-memory bank, the follow-up agent still wakes up blind.

That is the case for a shared memory layer instead of a per-tool one. Vilix AI is cloud-hosted, so there is no infrastructure to run: the voice pipeline, the scheduled agents, and the human's own AI tools all read and write the same memory over MCP. It stores full conversation history, not just extracted facts, so the actual exchange is revisitable. Retrieval is semantic, so recall finds what was meant rather than only what was typed, with keyword matching alongside for exact strings like order IDs and policy names.

Because the memory is shared, a correction only needs to happen once. Say "we no longer offer Saturday slots" to any connected tool and that becomes the truth for every agent reading the memory, voice or scheduled. The free plan is free forever, there is a 7-day Pro trial with no credit card, and everything can be exported or deleted at any time in a portable format.

The test that settles it

If you are running a Pipecat voice agent in production, run this test: call it, state a preference or a fact it should keep, hang up, call back from the same number, and see what happens. If it asks the same intake questions again, you know which of the four decisions above nobody made yet.

Your pipeline is fine. It just needs a memory that outlives the call. Start with the free plan and give your callers a voice agent that recognizes them the second time.

Get Started for Free

Persistent memory across ChatGPT, Claude, and the AI tools you already use in Vilix AI.

Get Started for Free

Free forever, no credit card.

Keep reading
Does LangGraph Remember Between Runs?

Does LangGraph Remember Between Runs? The short answer: only inside one thread. LangGraph's checkpointer saves your graph's state under a thread_id, so the same conversation can resume after a crash or a restart. A new thread_id starts with a blank slate. For memory that survives across threads and sessions, you need LangGraph's store, a separate system you have to wire in deliberately. What does a LangGraph checkpointer actually remember? A checkpointer snapshots your graph's state after ev

Your Scheduled Research Agent Rediscovers the Internet Every Morning. Here's the Memory Pattern.

Your Scheduled Research Agent Rediscovers the Internet Every Morning. Here's the Memory Pattern. Every Monday at 6 AM, a research agent wakes up and produces a market brief: funding rounds, product launches, pricing moves, the works. Every Monday the brief lands on time. And every Monday the agent builds it the same way: by rediscovering the entire internet from nothing. It starts with the landscape. Which companies are in this space, what they charge, who funds them. It researched all of this

Relay.app Is Gone, and So Is Everything It Remembered About Your Work

Relay.app Is Gone, and So Is Everything It Remembered About Your Work Relay.app shut down in September 2026. Free users lost access on August 15, paying customers on September 14, and then it was over: the shutdown notice went up on relay.app, and every account, workflow, and run history was permanently deleted. Founder Jacob Bank and part of the team moved to Google's Chrome team. Good product, talented team, and it still died, because standalone automation vendors keep getting absorbed by big