# Website

> Keep a website in a memory of its own, from its sitemap. A page that changed is read again and teaches only what changed, a page no longer listed is deleted with what it taught, and your assistant recalls from the site beside each user's own memory.

Keep a website in a memory of its own: your help centre, your prices, your opening hours. This sync reads
the site's sitemap, fetches each page and adds it under its address, in a space for the site. Run it on a
schedule. A page whose `lastmod` hasn't changed isn't fetched again, a page that changed teaches only what
changed, and a page the sitemap no longer lists is deleted with what it taught. Your assistant then recalls
from the site beside each user's own memory.

## Install

```bash
pip install httpx geniffy
```

```bash
uv add httpx geniffy
```

Set `GENIFFY_API_KEY` from **API keys** in the Geniffy app.

## The sync

```python
import xml.etree.ElementTree as ET
from urllib.parse import urlparse

import httpx
from geniffy import BadRequestError, Geniffy

geniffy = Geniffy()                               # reads GENIFFY_API_KEY
LABELS = {"channel": "website"}
SITEMAP = "{http://www.sitemaps.org/schemas/sitemap/0.9}"


def web() -> httpx.Client:
    return httpx.Client(timeout=30, follow_redirects=True, headers={"User-Agent": "site-sync/1.0"})


def listed(client: httpx.Client, sitemap_url: str) -> dict[str, str]:
    """Every page the sitemap lists, with when it last changed ("" when it doesn't say). A sitemap index is
    followed to each sitemap it names."""
    pages, todo = {}, [sitemap_url]
    while todo:
        root = ET.fromstring(client.get(todo.pop()).raise_for_status().content)
        for entry in root:
            loc = (entry.findtext(SITEMAP + "loc") or "").strip()
            if entry.tag == SITEMAP + "sitemap" and loc:
                todo.append(loc)
            elif entry.tag == SITEMAP + "url" and loc:
                pages[loc] = (entry.findtext(SITEMAP + "lastmod") or "").strip()
    return pages


def space_for(sitemap_url: str) -> str:
    return "site_" + urlparse(sitemap_url).hostname


def sync(sitemap_url: str, last: dict[str, str] | None = None) -> dict[str, str]:
    """Bring a website into its own space. Returns what the sitemap listed: pass it in next time, and a page
    whose lastmod hasn't changed isn't fetched again."""
    mem, last = geniffy.space(space_for(sitemap_url)), last or {}
    with web() as client:
        pages = listed(client, sitemap_url)
        for url, lastmod in pages.items():
            if lastmod and last.get(url) == lastmod:
                continue                          # unchanged since the last run
            got = client.get(url)
            if got.status_code != 200 or "html" not in got.headers.get("content-type", ""):
                continue                          # down for now, or not a page: memory keeps what it had
            try:
                mem.memories.add_file(got.content, filename="page.html", title=url, external_id=f"web:{url}",
                                      labels=LABELS)
            except BadRequestError:               # a page with no text in it
                pass
    mem.sources.delete_labelled(LABELS, keep={f"web:{url}" for url in pages})
    return pages
```

Run it on a schedule, and keep what each run returns for the next: `pages = sync(SITEMAP_URL, pages)`. Your
server fetches the pages itself, so a site behind a sign-in works too: give `web()` the cookie or header it
needs. Pages without a `lastmod` are fetched every run, and one that hasn't changed teaches nothing new.

## Recall from it

The site has a space of its own, so every user's conversation can draw on it, beside what is known about that
user:

```python
site = geniffy.space(space_for(SITEMAP_URL)).context(question)
user = geniffy.space(f"user_{user_id}").context(question)
system = f"<site>\n{site}\n</site>\n\n<memory>\n{user}\n</memory>"
```

## How it behaves

- **Each page is one source,** under its address and titled by it, so what recall gives back cites the page.
  Changed, only the paragraphs that changed are learned. See [Your own ids](https://docs.geniffy.com/add-memories/your-own-ids).
- **A page the sitemap no longer lists goes,** with what it taught. A page that is down for a run keeps what
  it had until it is back.
- **Only pages are read:** images, PDFs and anything else the sitemap lists that isn't HTML are left out. For
  PDFs, add them as [files](https://docs.geniffy.com/add-memories/files).
- **One space for each site,** named by its host, so two sites never mix. To forget a site, erase its
  space: `geniffy.forget_space(space_for(SITEMAP_URL))`.

Source: https://docs.geniffy.com/integrations/website
