ContraptionSoft solutions

NOTES · OCT 9, 2026

My inbox was a shit show, so I fixed it with Jev


A free skill for Claude Code or any agent: it designs buckets from your real mail, sorts your inbox with Jev, and keeps it sorted for a fraction of a cent per email.

By Tyler Malone

How it got this bad

I'll be honest, I let my email get bad. 3,204 threads sitting in the inbox, 1,626 of them unread.

I'd fixed it once before. I had an OpenClaw agent sorting my mail for a while, and it worked, but it cost too much. It talks too much, and you pay for every word when all you needed was a label. So I turned it off, and the inbox slid right back.

And yeah, I know Gmail has filters. They're not that good, and I'm not sitting there writing a rule for every sender who's ever emailed me.

There's a new way to do it now. Jev, baby.

Jev is a cheap model from TypeSafe that answers yes/no questions with a number instead of writing text. No essays, just a score.

It sorted every thread, archived 70% of them, and left me with 979. Every email got exactly one label: client email, bills, receipts, tool notifications, newsletters, account security, personal, plus a Review label for anything unclear. Jev's part of the bill was about 22 cents, at TypeSafe's listed price.[1] A job now sorts new mail every 15 minutes, so it stays that way.

I packaged the whole process as a skill. Give it to your AI agent and it walks you through the setup on your email provider and your machine.

Not into terminals? We'll set it up for you. Just mention the inbox skill.

What your agent will do

  1. Ask about your setup: your email provider, where a scheduled job can run, and whether other tools read your inbox.
  2. Back up your current labels, so everything can be undone.
  3. Read about 300 of your real emails and propose buckets built around what each email needs from you. You approve them.
  4. Dry run 200 emails without touching your mailbox, then fix the misses by rewording the questions.
  5. Sort the whole inbox. It archives only what it's confident about. Anything uncertain gets a label but stays in your inbox.
  6. Install a job that sorts new mail every 15 minutes.

It stops for your OK at every step that matters. The skill includes our working Gmail code as a reference, and the same steps carry over to Outlook, Fastmail, and IMAP.

What you'll need

  • An AI agent that can run commands, like Claude Code
  • A Jev API key from TypeSafe's console
  • A machine that's on most of the time, for the scheduled job
  • For Gmail: about 10 minutes in Google Cloud to create a sign-in client for your inbox. Your agent tells you exactly what to click.

Get the skill

For Claude Code on Mac or Linux:

mkdir -p ~/.claude/skills/inbox-triage && cd ~/.claude/skills/inbox-triage && \
curl -fsSLO https://contraptionsoft.com/skills/inbox-triage/SKILL.md && \
curl -fsSLO https://contraptionsoft.com/skills/inbox-triage/gmail_example.py

Then say: "set up inbox triage."

On Windows, or with any other agent, point it at https://contraptionsoft.com/skills/inbox-triage/SKILL.md and ask it to install the skill and follow the steps.

Here's exactly what you're installing:

View the files2 files
---name: inbox-triagedescription: Set up automatic email triage for the user. Design label buckets from a sample of their real mail, sort the whole inbox with Jev (a cheap model that scores yes/no questions), archive only the noise it's sure about, and keep new mail sorted on a schedule. Works with any provider that has an API (Gmail, Outlook/Microsoft Graph, Fastmail, IMAP) and any machine that can run a scheduled job. Use when the user wants to organize, label, triage, or clean up an inbox.--- # Inbox triage with Jev You're building three things for the user: 1. A one-time sort of their whole inbox, with one label (or folder) per conversation.2. Automatic archiving of the noise, but only when the scoring model is confident.3. A scheduled job that keeps new mail sorted. Work in the phases below. Stop for the user's OK at each **checkpoint**. Nothing touches their mailbox until phase 6. `gmail_example.py` in this folder is a complete, stdlib-only Python version for Gmail. Use it as a reference, ordirectly if they're on Gmail. Don't assume Gmail: adapt the same steps to whatever they use. ## 0. Learn their setup (ask; don't guess) - **Provider and access.**  - Gmail: a Google Cloud project with the Gmail API on and an OAuth client with the `gmail.modify` scope.  - Outlook / Microsoft 365: Microsoft Graph with `Mail.ReadWrite`.  - Fastmail: an API token.  - Anything else: IMAP with an app password.- **Where the scheduled job will run.** Options: an always-on server (systemd timer or cron), a Mac (launchd), Windows  (Task Scheduler), or a cloud function. A laptop works, but only sorts while it's awake.- **Whether other tools read this inbox** (a CRM, helpdesk, or lead parser). If so, the scheduled job holds back new  mail for a few minutes so those tools see it first.- **A Jev API key** from typesafe.ai.- **Secrets:** keep them in a file only the user can read, and never print them. ## 1. Back up before anything else Record every message's current labels or folders to a file. This is the undo record. On Gmail, you don't need tofetch every message: list the message IDs under each label and invert the result. That's a few dozen calls insteadof thousands. ## 2. Sample their real mail Pull about 300 inbox emails, spread across their existing labels and folders, with at least 8 from each group so smallones show up. Keep the sender, subject, and first ~500 characters of each. This step only reads. ## 3. Design the buckets (you do this) Read the sample and propose 6–12 buckets. Rules: - **Start from scratch.** Every email is about to be re-sorted, so their current labels are only hints. Renaming,  merging, splitting, or dropping labels costs nothing. Models tend to preserve existing structure; don't.- **Organize by what the mail needs from the user** (reply, pay, read later, ignore), not by who sent it.- **No catch-all bucket.** `Review` already exists for anything that fits nowhere.- Each bucket gets a `name`, `archive: true/false`, and a `question`.- **Questions are yes/no, about one email.** List the concrete kinds of email that belong, then say what doesn't.  Jev scores every question independently, so vague questions overlap. For example:  "Is this a receipt, order confirmation, or payment confirmation for something already paid, needing no action?  Not bills still due, failed payments, statements, shipping updates, or marketing."- Also write a one-line `owner` description (whose inbox this is and what they do). It goes in front of every email  sent to Jev. **Checkpoint:** show the user the buckets as a table (name, archive, what goes in it, example subjects from thesample). They will edit it. Save the config only after they approve. ## 4. How scoring works Send one Jev request per email, with all the questions in it. You get back a 0–1 probability for each: ```POST https://api.typesafe.ai/v1/systemoneAuthorization: Bearer $JEV_API_KEY{"model": "jev-latest", "state": "<owner line>\n\nFrom: ...\nSubject: ...\n\n<first ~1,500 chars of body>", "questions": {"receipts": {"type": "noul", "instructions": "<question>"}, ...}}-> {"answers": {"receipts": {"noul": 0.94}, ...}, "usage": {...}}``` The highest score wins. If nothing scores at least the threshold (start at 0.5), the email goes to `Review` and staysin the inbox. ## 5. Dry run (still no mailbox changes) Score 200 sampled emails. Write a CSV with: top bucket, score, second bucket, second score, sender, subject. - **Read every row yourself.** Show the user the low scores, the close calls, and any cluster of the same mistake.- **Fix mistakes by rewording questions, not with code.** Add the concrete kind of email that got missed. Rerun the  same 200 and compare. Two rounds is typical.- **A low score usually means no bucket fits.** That's what `Review` is for.- **A confident wrong answer usually means a bucket is missing.** Add the bucket; don't tune the threshold. **Checkpoint:** show the accuracy you found and the threshold you'd use, and get a go-ahead. ## 6. Sort the inbox - **Which message to judge:** for each conversation, the newest message the user didn't send. Their own reply says  nothing about what the thread is.- **One update per conversation:** add the bucket label, remove their old labels, and archive only when **all three**  are true:  - the bucket archives,  - the score is ≥ 0.7,  - it leads second place by ≥ 0.1.   Anything else gets its label but stays in the inbox. An extra email in the inbox costs a glance; a wrongly  archived one can cost a customer.- **Folder-based providers** (Outlook, IMAP): moving to the bucket's folder is the archive step. For "label but  keep in inbox", use a category or flag instead.- **Smoke test 5 conversations first.** Have the user check them in their own mail app (**checkpoint**).- **Then the full run:**  - run it in the background,  - record each finished conversation in a checkpoint file so a crash or rerun resumes where it stopped,  - make failures print which API refused and why. ## 7. Keep it sorted - **What the job does:** runs every ~15 minutes, finds inbox mail with **no bucket yet** (a search or filter, not a  stored "last seen" marker, so a failed run is just picked up next time), and sorts it the same way.- **New mail:** skip anything younger than ~10 minutes if other tools read the inbox.- **Replies:** conversations that already have a bucket keep it when a reply arrives.- **Install it** with whatever scheduler their machine uses, and confirm it has actually run once.- **Tell the user** how to see its log, pause it, and change a bucket (edit the question; the next run uses it). ## Gotchas - **The provider's rate limit sets the pace, not Jev.**  - Gmail allows 6,000 quota units per minute per user, per Google Cloud project. Reading a thread costs 40 units    and changing it costs 10, so the ceiling is ~120 threads a minute. Other tools on the same project share it.  - On a 403/429, wait out the full minute and retry. Microsoft Graph sends 429 with a `Retry-After` header; honor    it.- **Access tokens usually expire after an hour.** Refresh them during long runs.- **Google OAuth:**  - An external app still in "Testing" gets refresh tokens that expire after 7 days, so the scheduled job dies a    week later. On Google Workspace, set the app to **Internal**.  - A redirect URI has to match the OAuth client exactly, and new ones can take a few minutes to start working.- **Reused label names:** if a bucket has the same name as an old label, don't remove the label you're adding in  the same request. Gmail rejects that with a 400.- **Don't trust a high score when the right bucket doesn't exist.** Vendor support tickets are the classic case.## What to expect From the run this skill comes from: 3,204 inbox threads went to 979, with 70% archived. The full run took about75 minutes, almost all of it waiting on Gmail's rate limit. Jev input tokens cost about $0.22 at list price. Thedry run found roughly 93% clearly right after one round of rewording questions.
SKILL.md · 140 lines · 7.9 KBDownload
#!/usr/bin/env python3"""Gmail triage: Jev sorts every inbox thread into one bucket label and archives the noise.From ContraptionSoft (contraptionsoft.com/notes/triage-email-with-jev). Python 3.8+, no dependencies.Data lives in $GMAIL_TRIAGE_DIR (default ~/gmail-triage). Its .env needs GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET,JEV_API_KEY; `auth` adds GMAIL_REFRESH_TOKEN. Buckets live in categories.json: {"owner": "...", "review_threshold": 0.5, "buckets": [{"name", "question", "archive"}]}.  gmail_example.py auth     one-time Gmail sign-in (gmail.modify): prints the link; `auth '<url>'` saves the refresh token  gmail_example.py backup   every message's labels -> labels-before.json (the undo record)  gmail_example.py sample   ~300 inbox emails spread across your current labels -> sample.json + sample.txt (read-only)  gmail_example.py dryrun   Jev scores 200 sampled emails -> dryrun.csv (no Gmail calls)  gmail_example.py apply    sort every inbox thread for real (resumable; progress in applied.jsonl)  gmail_example.py run      timer job: sort new inbox threads (no bucket label yet, at least 10 min old)  gmail_example.py test     self-check"""import base64, collections, csv, html, json, os, random, re, sys, threading, time, urllib.error, urllib.parse, urllib.requestfrom concurrent.futures import ThreadPoolExecutorfrom pathlib import Path HERE = Path(os.environ.get("GMAIL_TRIAGE_DIR", Path.home() / "gmail-triage"))OWN_ENV = HERE / ".env"SCOPE = "https://www.googleapis.com/auth/gmail.modify"REDIRECT = "http://localhost:8765"  # nothing listens here: you copy the dead URL back into the terminalAPI = "https://gmail.googleapis.com/gmail/v1/users/me"BODY_CHARS = 1500  def env_file(path):    if not path.exists():        return {}    pairs = (l.split("=", 1) for l in path.read_text().splitlines() if "=" in l and not l.startswith("#"))    return {k.strip(): v.strip().strip('"') for k, v in pairs}  def env():    return env_file(OWN_ENV)  def http(method, url, headers=None, body=None, form=False, timeout=60):    data = None    if body is not None:        data = urllib.parse.urlencode(body).encode() if form else json.dumps(body).encode()        headers = {**(headers or {}), "Content-Type": "application/x-www-form-urlencoded" if form else "application/json"}    req = urllib.request.Request(url, data=data, headers=headers or {}, method=method)    with urllib.request.urlopen(req, timeout=timeout) as r:        return json.loads(r.read() or b"{}")  def google_token(e, **params):    return http("POST", "https://oauth2.googleapis.com/token",                body={"client_id": e["GOOGLE_CLIENT_ID"], "client_secret": e["GOOGLE_CLIENT_SECRET"], **params}, form=True)  def code_from(pasted):    """The ?code= out of the localhost URL the browser got stuck on (or a bare code)."""    q = urllib.parse.parse_qs(urllib.parse.urlparse(pasted.strip()).query)    return q["code"][0] if "code" in q else pasted.strip()  def auth(pasted=None):    e = env()    if not pasted:        url = "https://accounts.google.com/o/oauth2/v2/auth?" + urllib.parse.urlencode({            "client_id": e["GOOGLE_CLIENT_ID"], "redirect_uri": REDIRECT, "response_type": "code", "scope": SCOPE,            "access_type": "offline", "prompt": "consent"})        print(f"1. Open this and sign in with the Gmail account to sort:\n\n{url}\n")        print("2. The browser ends on a localhost page that won't load. Run: triage.py auth '<that whole URL>'")        return    t = google_token(e, code=code_from(pasted), grant_type="authorization_code", redirect_uri=REDIRECT)    if "refresh_token" not in t:        sys.exit("Google returned no refresh token. Revoke the app at myaccount.google.com/permissions and rerun.")    lines = [l for l in (OWN_ENV.read_text().splitlines() if OWN_ENV.exists() else []) if not l.startswith("GMAIL_REFRESH_TOKEN=")]    OWN_ENV.write_text("\n".join(lines + [f"GMAIL_REFRESH_TOKEN={t['refresh_token']}"]) + "\n")    OWN_ENV.chmod(0o600)    p = gmail(t["access_token"], "/profile")    print(f"Connected {p['emailAddress']}: {p['messagesTotal']} messages, {p['threadsTotal']} threads.")  _token = {"value": None, "expires": 0}  def access_token():    """Cached; refreshed 10 minutes before Google's 1-hour expiry so long runs keep going."""    if time.time() < _token["expires"]:        return _token["value"]    e = env()    if "GMAIL_REFRESH_TOKEN" not in e:        sys.exit("No Gmail token yet: run `gmail_example.py auth` first.")    t = google_token(e, refresh_token=e["GMAIL_REFRESH_TOKEN"], grant_type="refresh_token")    _token.update(value=t["access_token"], expires=time.time() + t.get("expires_in", 3600) - 600)    return _token["value"]  def gmail(at, path, body=None, **params):    """GET, or POST when there's a body. `at` is ignored in favor of the cached token (kept so old callers work)."""    for wait in (61, 61, 61, 61, 61, None):  # 403/429 = per-user per-minute quota: wait out the window        try:            return http("POST" if body is not None else "GET", f"{API}{path}?{urllib.parse.urlencode(params, doseq=True)}",                        {"Authorization": f"Bearer {access_token()}"}, body)        except urllib.error.HTTPError as e:            if e.code not in (403, 429) or wait is None:                raise            time.sleep(wait + random.random())  def invert(by_label):    """{label: [msg ids]} -> {msg id: [labels]}"""    out = {}    for label, ids in by_label.items():        for i in ids:            out.setdefault(i, []).append(label)    return out  def backup():    """List each label's message ids (a few dozen calls) instead of fetching every message (thousands)."""    HERE.mkdir(parents=True, exist_ok=True)    at = access_token()    labels = {l["id"]: l["name"] for l in gmail(at, "/labels")["labels"]}    by_label = {}    for lid, name in labels.items():        ids, page = [], None        while True:            r = gmail(at, "/messages", labelIds=lid, maxResults=500, includeSpamTrash="true", **({"pageToken": page} if page else {}))            ids += [m["id"] for m in r.get("messages", [])]            page = r.get("nextPageToken")            if not page:                break        by_label[lid] = ids        print(f"{name}: {len(ids)}", flush=True)    messages = invert(by_label)    out = HERE / "labels-before.json"    out.write_text(json.dumps({"history_id": gmail(at, "/profile")["historyId"], "labels": labels, "messages": messages}))    print(f"Backed up labels on {len(messages)} messages -> {out}")  def body_text(part):    """text/plain anywhere in the MIME tree, else tag-stripped HTML."""    def find(p, mime):        if p.get("mimeType") == mime and p.get("body", {}).get("data"):            return base64.urlsafe_b64decode(p["body"]["data"] + "==").decode("utf-8", "replace")        return next((t for c in p.get("parts", []) if (t := find(c, mime))), None)    text = find(part, "text/plain") or re.sub(r"<[^>]+>", " ", html.unescape(re.sub(r"(?is)<(style|script).*?</\1>", "", find(part, "text/html") or "")))    return re.sub(r"\s+", " ", text).strip()  def fetch(at, mid):    return parse(gmail(at, f"/messages/{mid}", format="full"))  def parse(m):    h = {x["name"].lower(): x["value"] for x in m["payload"].get("headers", [])}    return {"id": m["id"], "thread": m["threadId"], "labels": m.get("labelIds", []), "from": h.get("from", ""),            "subject": h.get("subject", ""), "date": h.get("date", ""), "body": body_text(m["payload"])[:BODY_CHARS]}  def group_of(label_ids, names):    """Which existing bucket a message sits in: its hand-made label, else its Gmail tab."""    user = [names[l] for l in label_ids if l.startswith("Label_")]    return user[0] if user else next((l for l in label_ids if l.startswith("CATEGORY_")), "none")  def stratify(groups, target, floor, rng):    """~target picks, proportional to group size, at least `floor` from each group so small ones show up."""    total = sum(len(v) for v in groups.values())    return [i for ids in groups.values() for i in rng.sample(ids, min(len(ids), max(floor, round(target * len(ids) / total))))]  def sample(n="300"):    """Fetch ~n inbox emails spread across current labels/tabs -> sample.json, and sample.txt for a human or LLM to read."""    b = json.loads((HERE / "labels-before.json").read_text())    groups = {}    for mid, lids in b["messages"].items():        if "INBOX" in lids:            groups.setdefault(group_of(lids, b["labels"]), []).append(mid)    with ThreadPoolExecutor(4) as pool:        picked = list(pool.map(lambda mid: fetch(None, mid), stratify(groups, int(n), 8, random.Random(1))))    (HERE / "sample.json").write_text(json.dumps(picked))    counts = "\n".join(f"- {g}: {len(v)}" for g, v in sorted(groups.items(), key=lambda kv: -len(kv[1])))    lines = "\n".join(f"#{k} [{group_of(m['labels'], b['labels'])}] From: {m['from']} | Subject: {m['subject']} | {m['body'][:300]}"                      for k, m in enumerate(picked))    (HERE / "sample.txt").write_text(f"Inbox: {sum(map(len, groups.values()))} emails. Current labels/tabs:\n{counts}\n\n{lines}\n")    print(f"{len(picked)} emails -> {HERE / 'sample.txt'}")  def email_state(m, owner):    return f"{owner}\n\nFrom: {m['from']}\nDate: {m['date']}\nSubject: {m['subject']}\n\n{m['body']}"  def jev(state, buckets):    """One Jev call scores every bucket's question 0-1 -> ({name: score}, usage)."""    for attempt in range(3):        try:            r = http("POST", "https://api.typesafe.ai/v1/systemone", {"Authorization": f"Bearer {env()['JEV_API_KEY']}"},                     {"model": "jev-latest", "state": state,                      "questions": {b["name"]: {"type": "noul", "instructions": b["question"]} for b in buckets}}, timeout=30)            return {b["name"]: r["answers"][b["name"]]["noul"] for b in buckets}, r.get("usage", {})        except (urllib.error.URLError, TimeoutError, KeyError):            if attempt == 2:                raise            time.sleep(2 + attempt * 3)  def pick(scores, threshold):    """Highest score wins; nothing at or above the threshold -> Review."""    top = max(scores, key=scores.get)    return top if scores[top] >= threshold else "Review"  def dryrun(n="200"):    """Score sampled emails with Jev -> dryrun.csv. No Gmail calls at all: reuses sample.json."""    cfg = json.loads((HERE / "categories.json").read_text())    b = json.loads((HERE / "labels-before.json").read_text())    pool_ = json.loads((HERE / "sample.json").read_text())    picked = random.Random(2).sample(pool_, min(int(n), len(pool_)))    started = time.time()    with ThreadPoolExecutor(8) as pool:        results = list(pool.map(lambda m: jev(email_state(m, cfg["owner"]), cfg["buckets"]), picked))    took = time.time() - started    rows = []    for m, (scores, _) in zip(picked, results):        (top, s1), (second, s2) = sorted(scores.items(), key=lambda kv: -kv[1])[:2]        rows.append({"top": top, "score": round(s1, 2), "second": second, "second_score": round(s2, 2),                     "old": group_of(m["labels"], b["labels"]), "from": m["from"][:60], "subject": m["subject"][:90], "id": m["id"]})    rows.sort(key=lambda r: (r["top"], -r["score"]))    with open(HERE / "dryrun.csv", "w", newline="") as f:        w = csv.DictWriter(f, fieldnames=list(rows[0]))        w.writeheader()        w.writerows(rows)    tok_in = sum(u.get("input_tokens", 0) for _, u in results)    tok_out = sum(u.get("output_tokens", 0) for _, u in results)    print(f"{len(rows)} emails x {len(cfg['buckets'])} questions in {took:.1f}s, {tok_in} tokens in / {tok_out} out\n")    print("top bucket (ignoring threshold):")    for name, c in collections.Counter(r["top"] for r in rows).most_common():        print(f"  {name:24} {c}")    tops = sorted(r["score"] for r in rows)    print("\ntop score percentiles:", {p: tops[int(p / 100 * (len(tops) - 1))] for p in (5, 10, 25, 50, 75)})    for t in (0.3, 0.4, 0.5, 0.6, 0.7):        print(f"  threshold {t}: {sum(s < t for s in tops)} to Review")  def archive_ok(scores, threshold=0.7, margin=0.1):    """Archive only when Jev is sure: high score AND a clear lead over second place. Else label but keep in inbox."""    first, second = sorted(scores.values(), reverse=True)[:2]    return first >= threshold and first - second >= margin  def latest_incoming(messages):    """The newest message you didn't send (your own reply says nothing about what the thread is); else the newest."""    theirs = [m for m in messages if "SENT" not in m.get("labelIds", [])]    return (theirs or messages)[-1]  def setup():    """Buckets + Gmail label ids (creating any missing labels). Shared by apply and run."""    cfg = json.loads((HERE / "categories.json").read_text())    ids = {l["name"]: l["id"] for l in gmail(None, "/labels")["labels"]}    for name in [b["name"] for b in cfg["buckets"]] + ["Review"]:        if name not in ids:            ids[name] = gmail(None, "/labels", {"name": name, "labelListVisibility": "labelShow", "messageListVisibility": "show"})["id"]    old = {lid for lid in json.loads((HERE / "labels-before.json").read_text())["labels"] if lid.startswith("Label_")}    return {"buckets": cfg["buckets"], "threshold": cfg["review_threshold"], "owner": cfg["owner"], "ids": ids, "old": old,            "archive": {b["name"] for b in cfg["buckets"] if b["archive"]}}  def sort_thread(ctx, tid, msgs):    """Jev picks the bucket; one modify adds it, strips old hand-made labels, archives when sure. -> result row."""    m = parse(latest_incoming(msgs))    scores, usage = jev(email_state(m, ctx["owner"]), ctx["buckets"])    bucket = pick(scores, ctx["threshold"])    (_, s1), (second, s2) = sorted(scores.items(), key=lambda kv: -kv[1])[:2]    archived = bucket in ctx["archive"] and archive_ok(scores)    add = ctx["ids"][bucket]    present = {l for x in msgs for l in x.get("labelIds", [])}    # A bucket that kept an old label's name reuses that label: never remove the one being added,    # Gmail 400s on "add and remove the same label".    try:        gmail(None, f"/threads/{tid}/modify",              {"addLabelIds": [add], "removeLabelIds": sorted((present & ctx["old"]) - {add}) + (["INBOX"] if archived else [])})    except urllib.error.HTTPError as e:  # say which API refused and why, not just "400"        raise RuntimeError(f"thread {tid}: {e.code} {e.read().decode()[:300]}") from e    row = {"thread": tid, "bucket": bucket, "score": round(s1, 3), "second": second, "second_score": round(s2, 3),           "archived": archived, "from": m["from"][:80], "subject": m["subject"][:120],           "tokens_in": usage.get("input_tokens", 0), "tokens_out": usage.get("output_tokens", 0)}    with open(HERE / "applied.jsonl", "a") as f:        f.write(json.dumps(row) + "\n")    return row  def inbox_threads(q=""):    out, page = [], None    while True:        r = gmail(None, "/threads", labelIds="INBOX", q=q, maxResults=500, **({"pageToken": page} if page else {}))        out += [t["id"] for t in r.get("threads", [])]        if not (page := r.get("nextPageToken")):            return out  def apply(limit=None):    """One-time: sort every inbox thread. Resumable via applied.jsonl."""    ctx = setup()    threads = inbox_threads()    done_file = HERE / "applied.jsonl"    done = {json.loads(l)["thread"] for l in done_file.read_text().splitlines()} if done_file.exists() else set()    todo = [t for t in threads if t not in done][:int(limit) if limit else None]    print(f"{len(threads)} inbox threads, {len(done)} already done, {len(todo)} to go", flush=True)    lock, started, n = threading.Lock(), time.time(), [0]     def one(tid):        sort_thread(ctx, tid, gmail(None, f"/threads/{tid}", format="full")["messages"])        with lock:            n[0] += 1            if n[0] % 100 == 0:                print(f"{n[0]}/{len(todo)} in {time.time() - started:.0f}s", flush=True)     with ThreadPoolExecutor(4) as pool:        for f in [pool.submit(one, t) for t in todo]:            if f.exception():                print("FAILED", repr(f.exception())[:300], flush=True)  # left out of applied.jsonl: a rerun retries it    rows = [json.loads(l) for l in done_file.read_text().splitlines()]    print(f"\ndone in {time.time() - started:.0f}s. {sum(r['archived'] for r in rows)} archived, "          f"{len(rows) - sum(r['archived'] for r in rows)} left in inbox", flush=True)    for name, c in collections.Counter(r["bucket"] for r in rows).most_common():        print(f"  {name:24} {c}", flush=True)  MIN_AGE = int(os.environ.get("GMAIL_TRIAGE_MIN_AGE", 600))  # seconds: lets other tools (CRM, helpdesk) see new mail before it can be archived  def unsorted_query(ctx):    """Inbox threads with no bucket label = new mail. No cursor to lose: a failed run is just picked up next time."""    return " ".join(f'-label:"{name}"' for name in [b["name"] for b in ctx["buckets"]] + ["Review"])  def run():    """Timer job: sort new inbox threads. Threads that already have a bucket keep it, even when a reply arrives."""    ctx = setup()    todo, now = inbox_threads(unsorted_query(ctx)), time.time()    for tid in todo:        msgs = gmail(None, f"/threads/{tid}", format="full")["messages"]        if int(latest_incoming(msgs)["internalDate"]) / 1000 > now - MIN_AGE:            print(f"skip {tid}: under {MIN_AGE // 60} min old, next run", flush=True)            continue        try:            r = sort_thread(ctx, tid, msgs)            print(f"{r['bucket']:22} {r['score']:.2f} {'archived' if r['archived'] else 'inbox   '} {r['from'][:40]} | {r['subject'][:70]}", flush=True)        except Exception as e:  # one bad thread shouldn't block the rest; it stays unsorted and is retried next run            print(f"FAILED {tid}: {e!r}"[:400], flush=True)  def test():    q = unsorted_query({"buckets": [{"name": "newsletters"}, {"name": "receipts"}]})    assert q == '-label:"newsletters" -label:"receipts" -label:"Review"'    assert archive_ok({"a": 0.9, "b": 0.5}) and not archive_ok({"a": 0.9, "b": 0.85}) and not archive_ok({"a": 0.65, "b": 0.1})    assert latest_incoming([{"id": 1}, {"id": 2, "labelIds": ["SENT"]}])["id"] == 1    assert latest_incoming([{"id": 1, "labelIds": ["SENT"]}])["id"] == 1    assert pick({"a": 0.9, "b": 0.2}, 0.5) == "a"    assert pick({"a": 0.4, "b": 0.2}, 0.5) == "Review"    assert group_of(["INBOX", "CATEGORY_UPDATES", "Label_5"], {"Label_5": "dev-noise"}) == "dev-noise"    assert group_of(["INBOX", "CATEGORY_UPDATES"], {}) == "CATEGORY_UPDATES"    picks = stratify({"big": list(range(1000)), "tiny": ["a", "b", "c"]}, 100, 8, random.Random(1))    assert len(picks) == 100 + 3 and {"a", "b", "c"} <= set(picks)    mime = {"mimeType": "multipart/alternative", "parts": [        {"mimeType": "text/html", "body": {"data": base64.urlsafe_b64encode(b"<style>x{}</style><p>Hi &amp; bye</p>").decode()}}]}    assert body_text(mime) == "Hi & bye"    assert code_from("http://localhost:8765/?code=4/0Ab-x_Y&scope=https://www.googleapis.com/auth/gmail.modify") == "4/0Ab-x_Y"    assert code_from("  4/0Ab-x_Y \n") == "4/0Ab-x_Y"    assert invert({"INBOX": ["a", "b"], "Label_5": ["b"]}) == {"a": ["INBOX"], "b": ["INBOX", "Label_5"]}    print("ok")  if __name__ == "__main__":    {"auth": auth, "backup": backup, "sample": sample, "dryrun": dryrun, "apply": apply, "run": run, "test": test}[sys.argv[1] if len(sys.argv) > 1 else "test"](*sys.argv[2:])
gmail_example.py · 373 lines · 19.2 KBDownload

What leaves your machine

To sort an email, the sender, subject, date, and first ~1,500 characters of the body go to TypeSafe's API to be scored. During setup, your agent also reads the 300-email sample, so that sample goes to whichever AI provider runs your agent. Your API keys, your label backup, and everything else stay on your machine. If that's a concern for your inbox, read TypeSafe's privacy policy first.[2]

How it works

Each bucket becomes one yes/no question for Jev, like "Is this a receipt for something already paid?", and the highest score wins. It can't invent a folder.

The score is what makes automatic archiving reasonable. The skill only archives when the top answer scores high and clearly beats the runner-up. Everything else keeps its label and stays in your inbox, and the backup means any mistake can be undone.

Your agent does the expensive thinking once, during setup, as a normal session. After that, every new email costs a fraction of a cent.

The slow part is your email provider's speed limit. Gmail lets an app read and relabel about 120 threads a minute,[3] and in practice it runs slower, so a few thousand threads takes an hour or so.


Want this set up on a shared team inbox, or wired into your CRM or helpdesk? That's what we do. See what we build or get in touch.

Sources

  1. TypeSafe AI, Jev
  2. TypeSafe AI, Privacy Policy
  3. Google, Gmail API usage limits