Pipecat
Give a Pipecat voice agent a memory of each caller. One frame processor sits in front of the LLM: on every turn it puts what is known about the caller into the system message, and when the call ends it saves the call. The next call starts knowing what was said in this one.
Install
pip install "pipecat-ai[anthropic]" geniffyuv add "pipecat-ai[anthropic]" geniffyAdd the speech and transport services your bot already uses. Set GENIFFY_API_KEY from API keys in the
Geniffy app.
The memory processor
from geniffy import AsyncGeniffy
from pipecat.frames.frames import EndFrame, Frame, LLMContextFrame
from pipecat.processors.frame_processor import FrameDirection, FrameProcessor
geniffy = AsyncGeniffy() # reads GENIFFY_API_KEY
def text_of(message: dict) -> str:
content = message.get("content")
if isinstance(content, str):
return content
parts = content or []
return " ".join(p.get("text", "") for p in parts if p.get("type") == "text")
class GeniffyMemory(FrameProcessor):
"""What is known about the caller before each reply; the call, saved at the end."""
def __init__(self, user_id: str, instructions: str):
super().__init__()
self.mem = geniffy.space(f"user_{user_id}")
self.instructions = instructions
self.context = None
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
if isinstance(frame, LLMContextFrame):
self.context = frame.context
await self.recall()
elif isinstance(frame, EndFrame) and self.context:
said = [m for m in self.context.get_messages() if isinstance(m, dict)]
await self.mem.memories.add(messages=said, title="Call")
await self.push_frame(frame, direction)
async def recall(self):
messages = [m for m in self.context.get_messages() if isinstance(m, dict)]
asked = next((m for m in reversed(messages) if m.get("role") == "user"), None)
question = text_of(asked) if asked else "What matters about this caller?"
known = await self.mem.context(question)
prompt = f"{self.instructions}\n\n<memory>\n{known}\n</memory>"
system = {"role": "system", "content": prompt}
rest = self.context.get_messages()
if rest and isinstance(rest[0], dict) and rest[0].get("role") == "system":
rest = rest[1:]
self.context.set_messages([system, *rest])The processor owns the system message: it rebuilds it on every turn from your instructions and what is known about the caller that bears on what they just said, each line with where it came from. When nothing is known, the block says so in one sentence, so the agent says it doesn't know instead of guessing.
In your pipeline
Put the processor between the user aggregator and the LLM, and start the context empty:
memory = GeniffyMemory(
user_id=caller.id,
instructions="You are a friendly voice assistant. Keep answers short.",
)
context = LLMContext()
aggregators = LLMContextAggregatorPair(context)
pipeline = Pipeline([
transport.input(),
stt,
aggregators.user(),
memory, # what is known about the caller, before the LLM
llm,
tts,
transport.output(),
aggregators.assistant(),
])The call is saved when an EndFrame passes through, so end calls with one rather than cancelling the task:
@transport.event_handler("on_client_disconnected")
async def on_client_disconnected(transport, client):
await task.queue_frames([EndFrame()])Each call is added as one conversation titled Call. Geniffy keeps who said what, so what the caller said
becomes a fact about them, and what your agent said stays the agent's.