Amazon S3
Keep the documents in an S3 bucket in a memory: a company's policies, its contracts, its handbooks. This sync lists the files under a prefix, downloads the ones that changed and adds each under its key, in the space you name. Run it on a schedule. A file whose ETag hasn't changed isn't downloaded again, a changed one teaches only what changed, and a file that is gone from the bucket is deleted with what it taught.
Install
pip install boto3 geniffyuv add boto3 geniffySet GENIFFY_API_KEY from API keys in the Geniffy app. boto3 finds AWS credentials the usual way: the
role your server runs with, or the environment. They need s3:ListBucket on the bucket and s3:GetObject on
the prefix.
The sync
import boto3
from geniffy import BadRequestError, Geniffy
geniffy = Geniffy() # reads GENIFFY_API_KEY
READS = (".pdf", ".docx", ".pptx", ".xlsx", ".txt", ".md", ".csv", ".html") # the files Geniffy reads
def s3():
return boto3.client("s3")
def sync(bucket: str, prefix: str, space: str, last: dict[str, str] | None = None) -> dict[str, str]:
"""Bring the files under a prefix of a bucket into a space. Returns each file's ETag: pass it in next time,
and a file whose ETag hasn't changed isn't downloaded again."""
mem, last, now = geniffy.space(space), last or {}, {}
labels = {"channel": "s3", "from": f"s3://{bucket}/{prefix}"} # so one prefix never deletes another's
client = s3()
for page in client.get_paginator("list_objects_v2").paginate(Bucket=bucket, Prefix=prefix):
for item in page.get("Contents", []):
key = item["Key"]
if not key.lower().endswith(READS) or item["Size"] > 25 * 1024 * 1024:
continue # an image, a video, or over 25 MB
now[key] = item["ETag"]
if last.get(key) == item["ETag"]:
continue # unchanged since the last run
data = client.get_object(Bucket=bucket, Key=key)["Body"].read()
try:
mem.memories.add_file(data, filename=key.rsplit("/", 1)[-1], title=key,
external_id=f"s3:{bucket}/{key}", labels=labels)
except BadRequestError: # empty, or a scan with no text in it
pass
mem.sources.delete_labelled(labels, keep={f"s3:{bucket}/{key}" for key in now})
return nowRun it on a schedule, and keep what each run returns for the next:
etags = sync("acme-docs", "policies/", "acme_policies", etags).
Recall from it
known = geniffy.space("acme_policies").context("How many days of leave do we get?")Each file is cited by its key, so an answer can point to the document it came from.
How it behaves
- Each file is one source, under its bucket and key. Changed, only the paragraphs that changed are learned, and what was removed is taken back. See Your own ids.
- Only what changed is downloaded. A file whose ETag is the one the last run saw is left as it is.
- A file that is gone goes from memory too, with what it taught. Each prefix carries its own label, so syncing one prefix never deletes another's files from the same space.
- What can't be read is left out: images and video, files over 25 MB, and scans with no text in them. See Files.