A Discord bot that shows a moderator how comparable cases were handled — before the sanction lands. It never decides. It shows the record and hands the call back.
Every large moderation bot is reactive and serves the server. Precedent serves the mod team. It replaces the right-click sanction with one that records a structured case, and uses that record to answer a question no tool currently answers: have we handled this kind of thing this way before?
The decisions below were settled in order, because most of them constrain the ones after. Each names what was chosen, why, and what was rejected — the rejections carry as much of the design as the choices.
The decision chain
01Observe, or own?
The bot owns the moderation workflow
Right-click a message → Moderate
Discord's audit log gives action type, moderator, target, and a free-text reason that in practice is empty or the word “spam”. It does not give the offending message, and it is purged after 45 days. You cannot build the claim “these two cases were alike” on that — you would be comparing timeout to timeout and calling it insight.
Owning the flow means the case record carries the message, the channel, the rule, the sanction and the mod. The adoption cost is real, and it gets paid by being faster than the native path: the modal opens with the rule pre-guessed and the usual sanction pre-filled.
Rejected — observe-only (data too thin to support the core claim); the hybrid nag-thread (trains mods to resent the bot, and partial data corrupts the comparison set invisibly).
02Compare along what axis?
The server's own rules, plus a grown sub-reason
Rule 4 › slur · folksonomy, not taxonomy
The output isn't a metric, it's an argument aimed at a mod who is about to disagree with it. That argument only lands in the vocabulary they already defend decisions in. “Rule 4 has gone 24h eleven times of thirteen” is checkable. “The embedding model clustered these” is dismissed in a second, and rightly — near-identical text deserves different sanctions depending on who said it and to whom.
Every server has a “Rule 1: be respectful” that absorbs 60% of actions and destroys the comparison set. The fix is a sub-reason typed once and offered as a dropdown thereafter — it grows in the team's own language and costs nothing at setup.
Rejected — a built-in universal taxonomy (foreign vocabulary; appeals cite rule numbers, not your categories); content embeddings as the primary axis (opaque, and unarguable in front of a mod).
03When does the finding surface?
Inline at action time, and a team digest — never a private post-mortem
Silent on agreement · card only on divergence
Action time is the only moment the information can change anything; every other timing turns the bot into a critic. The digest carries what isn't visible case-by-case — a rule whose sanctions have quietly doubled since January, actions that spike after 1am.
The after-the-fact DM is the one to cut. It arrives once the decision is public, so the mod can only eat it silently or visibly reverse themselves. It buys no better outcome and makes the bot feel like a supervisor sitting behind them — a perception that kills this product.
rule Rule 4 › slur prior first offence team 11 of 13 comparable → 24h mute youwarn
It shows precedent, never a verdict, and it never blocks. Diverging asks for one reason — offered as buttons drawn from the reasons that team has actually used, with a text escape. Those overrides are the real dataset: they encode what counts as a legitimate reason to diverge, which today lives only in senior mods' heads.
Constraint — Discord modals can't update mid-fill, so the card is an ephemeral follow-up after submit. That makes the precedent lookup a sub-second indexed query, not an LLM call.
04What makes two cases comparable?
Rule, sub-reason, and the offender's prior record
Buckets: 0 priors · 1–2 · 3+
Escalation is what kills this bot in week two. A first-timer and someone on their fourth strike should not receive the same sanction — if the comparison set doesn't know that, every correctly-escalated action reads as a divergence and mods learn to click through without reading.
Coarse buckets. Real escalation reasoning isn't finer than new / known / chronic.
Active priors only. The bot must mirror the server's expiry policy exactly. A bot that disagrees with a mod about the facts of the case loses all authority instantly.
Minimum-n, or stay silent. Conditioning fragments the data fast. Silence is cheap; a confidently wrong precedent claim costs the mod's trust permanently, and you only get it once.
Rejected — a learned model over case features. It can't explain itself, and decision 03 committed to an output whose whole job is to be auditable.
05What happens before there's data?
The team declares a policy grid at setup
Rule × prior bucket → expected sanction
Minimum-n means an honest auditor is silent for weeks — the highest-risk window in the product's life. The grid inverts that into the strongest thing here: filling it in is a conversation the mod team has never had in one room, and two mods will fill the same cell differently. That argument, on day one, before a single case is recorded, is worth more than the first month of findings.
Expiry gets written down as part of it. Almost no team has decided how long a warning counts against you; senior mods carry it privately and disagree by a factor of three without ever discovering it.
Underneath — a quiet audit-log backfill, used only to populate prior counts. Too thin to be precedent, perfectly good for “has this person been actioned before” — without it, your first nudge puts a chronic offender in the first-timer pool.
06Compare against policy, or practice?
Both — and their disagreement is the signal
Agreement → precedent. Conflict → “unsettled, your call.”
Enforcing practice looked right and is a trap. If the nudge enforces practice and mods follow it, practice gets more uniform, the nudge gets more confident, practice tightens further. The bot converts whatever the team happened to be doing the month it was installed into permanent law. Drift is laundered into precedent; a norm nobody voted for acquires the authority of one.
Pure policy is no better — it's a rulebook checker, and the grid was filled in during one sitting by people guessing. So the bot reports both and refuses to adjudicate. The digest escalates persistent gaps as amendment proposals, with the accumulated override reasons as evidence, and the team votes. Drift can only become precedent by passing through a human decision.
07Who sees mod-level data?
Each mod sees only their own position
Self view + team aggregates. No per-mod comparisons.
Nearly every genuine consistency problem is legible without naming anyone. “Rule 4 ranges from warn to 7d with no pattern” carries the same information as “Sam is soft on Rule 4”, and only one of them starts a fight. Self-view is the only channel where “you're drifting harsh” lands as information rather than an accusation.
Rejected — full team visibility (the first digest that embarrasses a senior mod ends the install that afternoon); lead-only (makes it a tool admins deploy onto mods, poisoning the framing). Note: this is a policy boundary, not a technical one.
08Scope
Public, multi-tenant, self-hosted by you
Chosen over single-server depth
This was taken against the recommendation, deliberately — the mod-team tooling gap is real and unfilled, and reach is the point. What it forces:
Setup becomes the central product problem. You won't be in the room for any install.
You become the data controller for strangers' moderation records. Privacy policy, deletion path and retention limit are launch requirements, not chores.
Verification gates you at 100 servers, and sharding becomes mandatory past 2,500.
Decision 07's boundary becomes a promise to people who cannot inspect your database.
Standing risk
A consistency auditor can't be dogfooded alone — with one mod there is nothing to compare. The product needs a plural, active mod team and several actions a day before any cell fills. Breadth without one such server tells you nothing about whether the core claim holds.
09How does setup survive self-serve?
Auto-draft the grid, then fill it lazily
Ingest the rules channel → one question at first use
A 40-cell form is where the 2am install dies. Auto-drafting kills the blank page — reviewing a wrong suggestion is far easier than authoring from nothing, and a visibly wrong cell provokes a correction where an empty one provokes abandonment.
The rest fills at the moment of use: the first time anyone actions Rule 4, one question — what should this normally get for a first offence? The grid then only ever grows for rules the server actually enforces, which is about six of the twenty they've written down.
10What's stored, and for how long?
Content on a short clock, case records long-lived
Strike expiry severs identity
The tension between prior counts wanting years and privacy wanting deletion dissolves once you notice the comparison consumes no message text at all — only rule, sub-reason, bucket, sanction, mod, timestamp. Content exists solely so a human can check whether the cases really were comparable, and that utility decays fast.
So the retention trigger is a decision already made: at strike expiry, sever the user link. The case survives forever as an anonymous data point about how this server treats Rule 4; it stops being personal data about a person. Erasure requests become free to honour, because you never needed the identity you're deleting.
On removal — purge after a 30-day grace period. Immediate deletion means one accidental kick during a permissions cleanup destroys a year of irreplaceable precedent; indefinite retention means holding moderation records for servers that fired you.
11What stops mods routing around it?
Make evasion visible, never costly
Surface it in the digest as a pattern
Two routes out. The cheap one is inventing a sub-reason — empty comparison set, minimum-n, silence. The cheaper one is the native right-click, which an interactions-only bot never sees at all. Both fail the same bad way: the data looks clean because the bypassed cases simply don't exist.
Since the bot never blocks anything, the defence has to be social. The digest reports “four new sub-reasons this month, all by one mod, all on cases that would have diverged” and “17% of actions bypassed the bot”. Nobody is blocked, nobody is accused, and the pattern sits at exactly the altitude decision 07 allows.
Rejected — locking the taxonomy. New sub-reasons are the mechanism that makes decision 02 work; gate them and the catch-all rule reabsorbs everything.
12Does any of this face members?
Own record yes, precedent never
/myrecord — strikes, sanctions, expiry
This bot manufactures discoverable evidence of inconsistency. “You muted me for 7 days and my mate got a warn” is currently a feeling; afterwards it's a database. Comparative data must stay internal, or a support tool becomes a weapon pointed at the people who installed it.
Own-record is pure upside — “what strikes do I have and when do they expire” is among the most common questions in any server and almost nothing answers it. It also blunts the grievance that fuels most appeals.
Be honest in the pitch — the data exists and admins can read it. A tool that makes a team more consistent also documents when they weren't. That trade is the actual deal on offer.
13Stack, hosting, sharding
TypeScript, discord.js, Postgres — gateway with sharding
Sharded from day one, because coverage is a metric
Every trigger here is an interaction, so this could have run as an HTTP endpoint with no gateway and no sharding, ever. It runs a gateway anyway for one reason: audit-log events are the only way to see the actions that bypassed the bot. Without them the coverage number in decision 11 doesn't exist, and blindness looks identical to health.
Stateless from the first commit. All state in Postgres, no in-memory cache keyed by guild. Do this and sharding stays a deployment concern instead of becoming a refactor — almost every painful sharding migration is really a caching migration.
ShardingManager with recommendedShards, processes on one box. Cross-machine needs a broker; that's a problem worth having later.
Lean intents — guilds and moderation events only. No message cache, no member cache.
Postgres, not SQLite. Every core operation is a relational aggregate scoped by guild.
~$40/month always-on with multiple shard processes. Per-guild load is a handful of interactions a day.
What ships
v1 is not a consistency auditor. It's a moderation tool that records structured cases. Auditing is what it grows into — and the policy grid is what makes the nudge useful on day one, with zero cases recorded, while the corpus builds underneath.
Most of it is sanction application, not auditing. Role hierarchy checks, missing permissions, the 28-day timeout cap, bans on users who already left, partial failures where the record saves but the mute doesn't. Owning the moderation flow means being more reliable than the native button — a bot that occasionally fails to apply a mute is uninstalled the same day, however good the auditing is.
Schema sketch
-- guild_id on every table from the first migration; never query without itguilds guild_id pk · strike_expiry_days · content_ttl_days
rules_channel_id · installed_at · purge_after
rules id · guild_id · ordinal · title · body
sub_reasons id · guild_id · rule_id · label · created_by · created_at
policy id · guild_id · rule_id · sub_reason_id? · prior_bucket
expected_sanction · expected_duration · set_by · set_at
cases id · guild_id · rule_id · sub_reason_id · prior_bucket
sanction · duration · moderator_id
subject_id? -- nulled at strike expiry
message_content? -- nulled at content_expires_at
source -- 'bot' | 'native' (coverage metric)
created_at · expires_at
overrides case_id · precedent_shown · reason_code · reason_text · created_at
Two nullable columns carry the whole retention design: subject_id dies at strike expiry, message_content at its TTL, and the case row survives both as anonymous precedent.
Verified against the docs
Both assumptions checked against Discord’s developer documentation on 29 Aug 2026. Both hold — and the check turned up a third thing that changes the install flow.
Confirmed
Message context menus deliver content without the privileged intent
The intent gates content, embeds, attachments, components and poll fields on message objects — with four stated exceptions. One of them is exactly this bot’s entry point:
“…the message a message context menu command is used on.”docs.discord.com/developers/events/gateway
The application-commands reference corroborates it: the example payload for a message command carries a full message object under data.resolved.messages, content included, with no intent prerequisite noted.
Confirmed
Audit-log events ride the GUILD_MODERATION intent
GUILD_AUDIT_LOG_ENTRY_CREATE is carried by GUILD_MODERATION (1 << 2), which is a standard intent, not a privileged one. The coverage metric in decision 11 is available without any special approval.
New consequence
It needs VIEW_AUDIT_LOG — and losing it fails silently
“This event is only sent to bots with the VIEW_AUDIT_LOG permission.”docs.discord.com/developers/events/gateway-events
That’s a permission the installing server grants, so it has to be in the invite scope — and a server can decline it or strip it later during a permissions tidy-up. If that happens, the bot receives no audit-log events and reports 0% of actions bypassed, which reads as perfect compliance. It is precisely the failure decision 11 was written to avoid: blindness that looks identical to health.
So v1 must probe for the permission rather than infer coverage from silence — and when it’s missing, the digest says the coverage figure is unavailable instead of printing a zero.
Net result
Precedent needs no privileged intents at all. That removes the review most likely to stall the 100-server verification gate, and it means the bot never holds message content it wasn’t explicitly handed by a moderator’s own click — a materially easier story to put in the privacy policy.
Identity
The mark is a ditto
A ditto mark means same as the one above. That is the bot's entire sentence, in two strokes. It also avoids the gavel and the scales, which claim an authority this design deliberately refuses — the tool never adjudicates, it shows the record.
app icondivergence16–32px
The second stroke tilting the other way is the divergence state — the same mark, one thing out of line. It reads at favicon size because it's two strokes and nothing else.
Palette
Drawn from carbon-copy forms and stamp ink rather than the courtroom: cool form-paper, near-black, and an aniline violet that reads as record rather than alarm. Note there is no red in the moderation flow at all — red is reserved for a sanction that genuinely failed to apply.
Form
#E6E9E3
Ground. Cool grey-green, not cream.
Ink
#171A19
Type, icon field.
Aniline
#57457F
Accent, stamp violet.
Concord
#3C6656
Matches precedent. Never celebratory.
Divergence
#8E6414
Attention, not alarm. The nudge.
Fault
#8A3A34
Errors only. Never a judgment.
Embed accents follow the same three states, so a mod learns the colour before they read the words: green means this matched, ochre means look at this, violet means policy and practice disagree and it's your call.
Type
Role
Face
Why
Display
Archivo 600/700
Official and form-like without being institutional cosplay.
Body
Source Serif 4
This is a document people read carefully and argue with.
Data
DM Mono
Case records, counts, labels. Typewriter-adjacent, tabular.
Still open
The dogfooding gap. No named server yet with a plural mod team and enough volume to fill cells. Everything above assumes one exists.
Minimum-n threshold. Not chosen. It's the dial between silence and being confidently wrong.
Name availability. A one-word English name on Discord is likely taken. Alternates held: Evenhand, Plumb, Praxis.
Grid auto-draft cost. An LLM call in the setup path, dependent on rules channels being parseable at all.