# GitHub

> Keep a repository's issues and pull requests in a memory of their own, so an assistant knows what was decided and why. Each thread is one source; a new comment teaches only what it adds, and only what was updated since the last run is fetched.

Keep a repository's issues and pull requests in a memory of their own, so a coding assistant or a support
bot knows what was decided and why. This sync adds each issue and pull request as one source, with every
person's comment on it, under its number and the label `channel: github`, in a space for the repository.
The first run reads everything; after that, only what was updated since the last run is fetched, and a new
comment teaches only what it adds. Comments from bots, and pull requests a bot opened, stay out.

## Install

```bash
pip install httpx geniffy
```

```bash
uv add httpx geniffy
```

Set `GENIFFY_API_KEY` from **API keys** in the Geniffy app. From GitHub, use a fine-grained token with read
access to the repository's issues and pull requests, or your GitHub App's installation token.

## The sync

```python
import httpx
from geniffy import Geniffy

geniffy = Geniffy()                               # reads GENIFFY_API_KEY
LABELS = {"channel": "github"}


def github(token: str) -> httpx.Client:
    return httpx.Client(base_url="https://api.github.com", timeout=30,
                        headers={"Authorization": f"Bearer {token}", "Accept": "application/vnd.github+json",
                                 "X-GitHub-Api-Version": "2022-11-28"})


def every(api: httpx.Client, url: str, params: dict | None):
    """Each item of a list GitHub pages, following its Link header."""
    while url:
        got = api.get(url, params=params).raise_for_status()
        yield from got.json()
        url, params = got.links.get("next", {}).get("url"), None    # the next link carries its own query


def space_for(repo: str) -> str:
    return "github_" + repo.replace("/", "_")


def thread(api: httpx.Client, repo: str, issue: dict) -> str:
    """An issue or pull request as one note: what was asked, and each person's comment after it."""
    kind = "Pull request" if "pull_request" in issue else "Issue"
    state = "merged" if (issue.get("pull_request") or {}).get("merged_at") else issue["state"]
    parts = [f"{kind} #{issue['number']}, {state}.",
             f"From {issue['user']['login']}, {issue['created_at'][:10]}:\n{issue.get('body') or issue['title']}"]
    for comment in every(api, f"/repos/{repo}/issues/{issue['number']}/comments", {"per_page": 100}):
        if comment["user"]["type"] != "Bot" and (comment.get("body") or "").strip():
            parts.append(f"From {comment['user']['login']}, {comment['created_at'][:10]}:\n{comment['body']}")
    return "\n\n".join(parts)


def sync(repo: str, token: str, since: str | None = None) -> str | None:
    """Bring a repository's issues and pull requests into its own space. Returns the time to pass next time:
    pass None the first time and everything is read; after that, only what was updated since is fetched."""
    mem, newest, seen = geniffy.space(space_for(repo)), since, set()
    params = {"state": "all", "sort": "updated", "direction": "asc", "per_page": 100}
    if since:
        params["since"] = since
    with github(token) as api:
        for issue in every(api, f"/repos/{repo}/issues", params):
            newest = max(newest or "", issue["updated_at"])
            if issue["user"]["type"] == "Bot":    # a dependency bump, a release bot
                continue
            ref = f"github:{repo}#{issue['number']}"
            mem.memories.add(thread(api, repo, issue), title=f"#{issue['number']} {issue['title']}",
                             said_at=issue["updated_at"], external_id=ref, labels=LABELS)
            seen.add(ref)
    if since is None:                             # everything was read: what is no longer there goes
        mem.sources.delete_labelled(LABELS, keep=seen)
    return newest
```

Run it on a schedule, and keep the time each run returns for the next: `since = sync("acme/api", token, since)`.
Now and then, pass `None` to read everything again: an issue moved to another repository, or deleted, goes
from memory then.

## Recall from it

```python
known = geniffy.space(space_for("acme/api")).context("Why did we move off Redis?")
```

Each issue and pull request is cited by its number and title, so an answer can point to the thread.

## How it behaves

- **Each issue and pull request is one source,** under its number, dated by its last update. A new comment
  sends the thread again, and only what it adds is learned. See [Your own ids](https://docs.geniffy.com/add-memories/your-own-ids).
- **Only what was updated is fetched.** After the first run, GitHub's `since` names the threads that changed.
- **People, not bots.** Comments from bots, and pull requests a bot opened, stay out.
- **Open, closed or merged** heads each thread, so an answer can tell a decision that shipped from one that
  was dropped.
- **One space for each repository,** so two repositories never mix. To forget one, erase its space:
  `geniffy.forget_space(space_for("acme/api"))`.

Source: https://docs.geniffy.com/integrations/github
