OneDrive
Keep each of your users' OneDrive in their memory. Your app already holds an access token for each user,
from their Microsoft sign-in with the Files.Read permission; this sync adds their Word, PowerPoint, Excel,
PDF and text files, each under its own id with the label channel: onedrive. It reads through Microsoft
Graph's delta query: the first run lists everything, and each run after it lists only what changed since,
so a run that finds nothing new downloads nothing. A changed file teaches only what changed, a file that is
gone takes what it taught with it, and when the user disconnects OneDrive, one call forgets everything that
came from it.
Install
pip install httpx geniffyuv add httpx geniffySet GENIFFY_API_KEY from API keys in the Geniffy app.
The sync
import httpx
from geniffy import BadRequestError, Geniffy, NotFoundError
geniffy = Geniffy() # reads GENIFFY_API_KEY
LABELS = {"channel": "onedrive"}
DRIVE = "/me/drive" # the user's own OneDrive
READS = (".pdf", ".docx", ".pptx", ".xlsx", ".txt", ".md", ".csv", ".html") # the files Geniffy reads
def graph(token: str) -> httpx.Client:
"""One user's OneDrive, with the access token from their Microsoft sign-in."""
return httpx.Client(base_url="https://graph.microsoft.com/v1.0", timeout=60, follow_redirects=True,
headers={"Authorization": f"Bearer {token}"})
def keep(mem, api: httpx.Client, item: dict) -> bool:
"""Add one file under its OneDrive id. False when it holds nothing Geniffy can read."""
if not item["name"].lower().endswith(READS) or item.get("size", 0) > 25 * 1024 * 1024:
return False # an image, a video, or over 25 MB
data = api.get(f"{DRIVE}/items/{item['id']}/content").raise_for_status().content
try:
mem.memories.add_file(data, filename=item["name"], external_id=f"onedrive:{item['id']}", labels=LABELS)
except BadRequestError: # empty, or a scan with no text in it
return False
return True
def forget(mem, item_id: str) -> None:
try:
mem.sources.delete(external_id=f"onedrive:{item_id}")
except NotFoundError: # never added: an image, a video
pass
def sync(user_id: str, token: str, delta_link: str | None = None) -> str:
"""Bring one user's OneDrive into their memory. Returns the link to pass next time: pass None the first
time and everything is read; after that, only what changed since the last run is fetched."""
mem = geniffy.space(f"user_{user_id}")
seen, link = set(), delta_link or f"{DRIVE}/root/delta"
with graph(token) as api:
while True:
got = api.get(link)
if got.status_code == 410 and delta_link: # too old to ask for: read everything again
return sync(user_id, token)
out = got.raise_for_status().json()
for item in out["value"]:
if "deleted" in item:
forget(mem, item["id"])
elif "file" in item: # not a folder, nor the drive itself
if keep(mem, api, item):
seen.add(f"onedrive:{item['id']}")
elif delta_link:
forget(mem, item["id"])
if "@odata.deltaLink" in out:
break
link = out["@odata.nextLink"]
if delta_link is None: # everything was read: what is no longer there goes
mem.sources.delete_labelled(LABELS, keep=seen)
return out["@odata.deltaLink"]
def disconnect(user_id: str) -> int:
"""The user disconnected OneDrive: everything that came from it goes. Returns how many files."""
return geniffy.space(f"user_{user_id}").sources.delete_labelled(LABELS)Run it on a schedule, and keep the link each run returns with the user for the next run. Microsoft's access tokens last about an hour, so refresh the user's before each run, as your Microsoft sign-in library does.
For a SharePoint document library instead, set DRIVE to /sites/{site-id}/drive, sign in with
Sites.Read.All, and keep the library in one space of its own rather than one per user.
How it behaves
- Each file is one source, under its OneDrive id. Changed, only the paragraphs that changed are learned, and what was removed is taken back. See Your own ids.
- Only what changed is fetched. After the first run, the delta link names the files that changed, and a run with nothing new downloads nothing. When Microsoft asks for a fresh start (410 Gone), the run reads everything again and puts memory right.
- A file that is gone goes from memory too, with what it taught.
- What can't be read is left out: folders, images and video, files over 25 MB, and scans with no text in them. See Files.
- Recall can keep to OneDrive, or leave it out:
mem.context(question, labels={"channel": "onedrive"}). See Labels. - Disconnecting forgets it all in one call, and nothing the user added another way.