Confluence
Keep a Confluence space in a memory of its own: your team's decisions, runbooks and how-tos, so an assistant
answers from your own pages and names the page each answer came from. This sync adds each page under its id
with the label channel: confluence, in a space for the Confluence space. The first run reads every page;
after that, only the pages changed since the last run are read, and a changed page teaches only what changed.
A page moved to the trash is deleted from memory with what it taught.
Install
pip install httpx geniffyuv add httpx geniffySet GENIFFY_API_KEY from API keys in the Geniffy app. From Atlassian, make an API token for an account
that can read the space, and use it with that account's email address.
The sync
import httpx
from geniffy import BadRequestError, Geniffy, NotFoundError
geniffy = Geniffy() # reads GENIFFY_API_KEY
LABELS = {"channel": "confluence"}
def confluence(site: str, email: str, api_token: str) -> httpx.Client:
"""Your Atlassian site, such as acme for acme.atlassian.net."""
return httpx.Client(base_url=f"https://{site}.atlassian.net", timeout=60, auth=(email, api_token))
def space_for(space_key: str) -> str:
return "confluence_" + space_key.lower()
def pages(api: httpx.Client, space_id: str, status: str, since: str | None = None):
"""The space's pages with this status, the most recently changed first, down to since."""
url, params = f"/wiki/api/v2/spaces/{space_id}/pages", {"status": status, "sort": "-modified-date",
"body-format": "storage", "limit": 100}
while url:
out = api.get(url, params=params).raise_for_status().json()
for page in out["results"]:
if since and page["version"]["createdAt"] <= since:
return # older than the last run: the rest are too
yield page
url, params = out.get("_links", {}).get("next"), None # the next link carries its own query
def sync(site: str, email: str, api_token: str, space_key: str, since: str | None = None) -> str | None:
"""Bring a Confluence space into its own memory. Returns the time to pass next time: pass None the first
time and every page is read; after that, only the pages changed since."""
mem, newest, seen = geniffy.space(space_for(space_key)), since, set()
with confluence(site, email, api_token) as api:
found = api.get("/wiki/api/v2/spaces", params={"keys": space_key}).raise_for_status().json()["results"]
space_id = found[0]["id"]
for page in pages(api, space_id, "current", since):
newest, ref = max(newest or "", page["version"]["createdAt"]), f"confluence:{page['id']}"
try:
mem.memories.add_file(page["body"]["storage"]["value"].encode(), filename="page.html",
title=page["title"], external_id=ref, labels=LABELS)
seen.add(ref)
except BadRequestError: # a page with no text in it
pass
for page in pages(api, space_id, "trashed"): # the trash is small: every run reads all of it
try:
mem.sources.delete(external_id=f"confluence:{page['id']}")
except NotFoundError:
pass
if since is None: # everything was read: what is no longer there goes
mem.sources.delete_labelled(LABELS, keep=seen)
return newestRun it on a schedule, and keep the time each run returns for the next:
since = sync("acme", "you@acme.com", token, "ENG", since). Now and then, pass None to read everything
again: a page deleted for good, or archived, goes from memory then.
Recall from it
known = geniffy.space(space_for("ENG")).context("How do we roll back a release?")Each page is cited by its title, so an answer can point to the page it came from.
How it behaves
- Each page is one source, under its Confluence id, so it stays one source when it is renamed or moved. Changed, only the paragraphs that changed are learned. See Your own ids.
- Only what changed is read. Pages come back most recently changed first, and a run stops at the first page older than the last run.
- Trashed goes. A page in the trash is deleted from memory with what it taught. Archived pages and pages deleted for good go when you next read everything.
- One space for each Confluence space, so two never mix. To forget one, erase its space:
geniffy.forget_space(space_for("ENG")).