# Amazon S3

> Keep the documents under a prefix of an S3 bucket in a memory, such as a company's policies. A file whose ETag hasn't changed isn't downloaded again, a changed one teaches only what changed, and a file that is gone is deleted with what it taught.

Keep the documents in an S3 bucket in a memory: a company's policies, its contracts, its handbooks. This
sync lists the files under a prefix, downloads the ones that changed and adds each under its key, in the
space you name. Run it on a schedule. A file whose ETag hasn't changed isn't downloaded again, a changed one
teaches only what changed, and a file that is gone from the bucket is deleted with what it taught.

## Install

```bash
pip install boto3 geniffy
```

```bash
uv add boto3 geniffy
```

Set `GENIFFY_API_KEY` from **API keys** in the Geniffy app. boto3 finds AWS credentials the usual way: the
role your server runs with, or the environment. They need `s3:ListBucket` on the bucket and `s3:GetObject` on
the prefix.

## The sync

```python
import boto3
from geniffy import BadRequestError, Geniffy

geniffy = Geniffy()                               # reads GENIFFY_API_KEY
READS = (".pdf", ".docx", ".pptx", ".xlsx", ".txt", ".md", ".csv", ".html")    # the files Geniffy reads


def s3():
    return boto3.client("s3")


def sync(bucket: str, prefix: str, space: str, last: dict[str, str] | None = None) -> dict[str, str]:
    """Bring the files under a prefix of a bucket into a space. Returns each file's ETag: pass it in next time,
    and a file whose ETag hasn't changed isn't downloaded again."""
    mem, last, now = geniffy.space(space), last or {}, {}
    labels = {"channel": "s3", "from": f"s3://{bucket}/{prefix}"}    # so one prefix never deletes another's
    client = s3()
    for page in client.get_paginator("list_objects_v2").paginate(Bucket=bucket, Prefix=prefix):
        for item in page.get("Contents", []):
            key = item["Key"]
            if not key.lower().endswith(READS) or item["Size"] > 25 * 1024 * 1024:
                continue                          # an image, a video, or over 25 MB
            now[key] = item["ETag"]
            if last.get(key) == item["ETag"]:
                continue                          # unchanged since the last run
            data = client.get_object(Bucket=bucket, Key=key)["Body"].read()
            try:
                mem.memories.add_file(data, filename=key.rsplit("/", 1)[-1], title=key,
                                      external_id=f"s3:{bucket}/{key}", labels=labels)
            except BadRequestError:               # empty, or a scan with no text in it
                pass
    mem.sources.delete_labelled(labels, keep={f"s3:{bucket}/{key}" for key in now})
    return now
```

Run it on a schedule, and keep what each run returns for the next:
`etags = sync("acme-docs", "policies/", "acme_policies", etags)`.

## Recall from it

```python
known = geniffy.space("acme_policies").context("How many days of leave do we get?")
```

Each file is cited by its key, so an answer can point to the document it came from.

## How it behaves

- **Each file is one source,** under its bucket and key. Changed, only the paragraphs that changed are
  learned, and what was removed is taken back. See [Your own ids](https://docs.geniffy.com/add-memories/your-own-ids).
- **Only what changed is downloaded.** A file whose ETag is the one the last run saw is left as it is.
- **A file that is gone goes from memory too,** with what it taught. Each prefix carries its own label, so
  syncing one prefix never deletes another's files from the same space.
- **What can't be read is left out:** images and video, files over 25 MB, and scans with no text in them.
  See [Files](https://docs.geniffy.com/add-memories/files).

Source: https://docs.geniffy.com/integrations/s3
