# LlamaIndex

> Give a LlamaIndex agent a memory of each of your users, with a memory block that recalls what is known, or a recall tool the agent calls.

Give a LlamaIndex agent a memory of each of your users. A memory block brings what is known about the user
into the agent's memory before it answers, or a recall tool lets the agent look things up itself; either
way, each exchange is saved after the run.

## Install

```bash
pip install llama-index-core llama-index-llms-anthropic geniffy
```

```bash
uv add llama-index-core llama-index-llms-anthropic geniffy
```

Set `ANTHROPIC_API_KEY`, and `GENIFFY_API_KEY` from **API keys** in the Geniffy app. Any LLM LlamaIndex
supports works the same way. These examples use `AsyncGeniffy`, the same client with every method awaited.

## Remember each user

```python
from geniffy import AsyncGeniffy
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.llms import ChatMessage
from llama_index.core.memory import BaseMemoryBlock, Memory
from llama_index.llms.anthropic import Anthropic

geniffy = AsyncGeniffy()                 # reads GENIFFY_API_KEY
agent = FunctionAgent(llm=Anthropic(model="claude-opus-5-5"),
                      system_prompt="You are a helpful assistant.")


class GeniffyBlock(BaseMemoryBlock[str]):
    """What is known about one of your users that bears on their latest message."""

    name: str = "geniffy"
    user_id: str

    async def _aget(self, messages: list[ChatMessage] | None = None, **kwargs) -> str:
        asked = next(m for m in reversed(messages or []) if m.role == "user")
        return await geniffy.space(f"user_{self.user_id}").context(asked.content or "")

    async def _aput(self, messages: list[ChatMessage]) -> None:
        pass                             # every exchange is saved after the run instead


async def chat(user_id: str, message: str) -> str:
    memory = Memory.from_defaults(
        session_id=f"user_{user_id}", memory_blocks=[GeniffyBlock(user_id=user_id)],
    )
    reply = str(await agent.run(user_msg=message, memory=memory))
    await geniffy.space(f"user_{user_id}").memories.add(messages=[
        {"role": "user", "content": message},
        {"role": "assistant", "content": reply},
    ])
    return reply
```

Call it with the user from your own sign-in:

```python
reply = await chat(user.id, "Who signs the Lumen renewal?")
```

LlamaIndex puts the block into the system message, so the agent sees what is known about this user, each
line with where it came from. When nothing is known, the block says so in one sentence, so the agent says it
doesn't know instead of guessing. The block saves nothing itself: LlamaIndex only hands a block the messages
that overflow its short-term memory, and the exchange should reach Geniffy every time.

## Let the agent look things up

To let the agent decide when to look something up, give it a `recall` tool bound to the user. The model
never sees or chooses whose memory it reads.

```python
from geniffy import AsyncGeniffy
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.anthropic import Anthropic

geniffy = AsyncGeniffy()
llm = Anthropic(model="claude-opus-5-5")
INSTRUCTIONS = ("You are a helpful assistant. Use recall before answering anything "
                "that depends on what the user said before.")


def agent_for(user_id: str) -> FunctionAgent:
    mem = geniffy.space(f"user_{user_id}")

    async def recall(query: str) -> str:
        """Look up what is known about the user, with where it came from."""
        return await mem.context(query)

    return FunctionAgent(llm=llm, tools=[recall], system_prompt=INSTRUCTIONS)


async def chat(user_id: str, message: str) -> str:
    reply = str(await agent_for(user_id).run(user_msg=message))
    await geniffy.space(f"user_{user_id}").memories.add(messages=[
        {"role": "user", "content": message},
        {"role": "assistant", "content": reply},
    ])
    return reply
```

The tool answers with `context()`, so the agent reads the same lines, with their sources, that the memory
block would hold, and the same sentence when nothing is known.

Source: https://docs.geniffy.com/integrations/llamaindex
