Content moderation you can audit.
A public, consensus-verified moderation registry built on GenLayer Intelligent Contracts.
Live demo | Public audit log | Deployment evidence | Roadmap
When a post is removed, you usually cannot see who decided, why, or whether anyone could check it. Moderation today is one company's backend, one model and one unaccountable decision, with no trail a third party can verify.
Modeq moves the decision itself onto GenLayer and makes the result public.
- Independent classification. A submission is classified by an LLM that runs independently on a set of five GenLayer validators. The result only lands on-chain if a majority of them agree on the decision-relevant fields.
- The model never gets the final call. It returns a structured classification (categories plus a confidence score). Fixed thresholds written in the contract turn that into ALLOW, FLAG or BLOCK. A model can misclassify, but it cannot talk its way past the threshold logic with a clever free-text answer.
- A permanent public record. Every case is stored on-chain and listable by anyone, with no wallet, account or rate limit. A forum or DAO can point at an append-only moderation log instead of saying "trust us".
flowchart LR
A["Submitter<br/>submit_content(text)"] --> B["Leader validator<br/>LLM classifies"]
B --> C{"Validators re-run<br/>the classification"}
C -- "decision and primary_category match" --> D["Case written on-chain<br/>status ACCEPTED"]
C -- "disagreement" --> E["No consensus<br/>nothing is stored"]
D --> F["Public audit log<br/>list_cases / get_case"]
The thresholds are constants in the contract, not model output.
| Condition | Verdict |
|---|---|
primary_category is none |
ALLOW |
| confidence above 70% | BLOCK |
| confidence above 40% | FLAG |
| otherwise | ALLOW |
Categories: spam, hate_speech, harassment, nsfw, violence, none.
The model's JSON is validated in Python before anything is stored. Unknown keys, a
category outside the allow-list, or a non-numeric or out-of-range confidence raise a
UserError. A leader run that fails this way is not accepted, and a validator that
cannot reproduce valid output votes against the leader instead of storing bad data.
![]() |
![]() |
| /app - submit text from a wallet and follow the transaction through consensus | /audit - every case ever made, filterable, with the rule that fired |
A write is not a spinner. /app follows the transaction through the wallet, the send, the
validators' consensus and the read-back, using the status the chain reports. The capture
below is a real submission on studionet: three of the five validators agreed, which is
enough for consensus, and the app says exactly that.
contracts/moderation_registry.py
| Method | Kind | Description |
|---|---|---|
submit_content(text) |
write | Classify and store a new case, returns its case_id |
get_case(case_id) |
view | One case: submitter, text, categories, primary category, confidence, decision, timestamp |
list_cases(offset, limit) |
view | Paginated audit log |
total_cases() |
view | Number of stored cases |
The consensus logic uses a custom gl.vm.run_nondet_unsafe(leader_fn, validator_fn) pair
instead of strict_eq. LLM output is never byte-identical across validators, so the
validator re-runs the classification and compares only decision and primary_category.
Confidence jitter such as 97% against 99% does not break consensus. The history of that
change, and the real bug it fixed, is in deploy/NOTES.md.
| Network | Contract | Cases |
|---|---|---|
| GenLayer Studio (studionet), used by the live app | 0xBC9b8c99889fe33f7650FA3530387Ee931AbD107 |
live on the audit page |
| Asimov testnet | 0xF95A5969c79706C7f4274D4e633315bD014C56Eb |
1 |
| Bradbury testnet | 0x13bfD75B34d2C106EA472F105811194352c30461 |
1 |
Case counts come from total_cases() on each contract. Reproduction steps are in
deploy/NOTES.md.
Four manual attacks were submitted through the live app against the studionet contract, each trying to talk the classifier into calling obvious spam clean. All four were classified as spam and blocked, and they remain visible in the audit log.
| Case | Attack | Result |
|---|---|---|
| 10 | Prompt injection | BLOCK, spam, 99% |
| 11 | Smuggling a literal decision: ALLOW JSON blob |
BLOCK, spam, 96% |
| 12 | Delimiter injection that closes the <content> wrapper |
BLOCK, spam, 95% |
| 13 | Fake "debug mode" jailbreak | BLOCK, spam, 98% |
In all four the model's own judgment held, so the strict shape validation was never
triggered by a live model. It is verified against a mocked non-compliant response by
test_decision_is_computed_not_trusted_from_model. Both layers are described honestly in
deploy/NOTES.md.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
genvm-lint check contracts/moderation_registry.py
pytest tests/direct/ -v # 8 tests, in memory, mocked LLM
gltest --network studionet tests/integration/ -v -s # real consensus on GenLayer StudioThe direct tests cover the three thresholds, the none override, rejection of malformed
output and unknown categories, a model trying to smuggle its own decision field, and
the listing methods. genvm-lint reports one expected warning: time.time() is called
inside the non-deterministic leader function on purpose, to stamp each case.
cd frontend
cp .env.example .env
npm install
npm run dev # http://localhost:3000
npm run build # production build
npm test # transaction tracker testsConnect MetaMask on /app. The app switches the wallet to the GenLayer network before
every write, and does not rely on the client library for that check. No funds are needed:
Studio is a free test network, and a brand-new empty account can submit. If a wallet shows
"Fee is not set" and keeps Approve disabled, open its fee settings and enter any small
custom fee such as 1 gwei; nothing is charged. Details are in
frontend/README.md.
contracts/ Intelligent Contract (Python)
tests/direct/ fast in-memory tests with a mocked LLM
tests/integration/ end-to-end test against GenLayer Studio
deploy/NOTES.md deployment steps, on-chain evidence, bugs found and fixed
frontend/ Next.js 16 app: landing, /app tool, /audit log
docs/screenshots/ images used in this README
- Multimodal moderation: accept an image alongside or instead of text through
gl.nondet.exec_prompt(images=[...]), and extend the category schema. - Per-community configurable thresholds and category sets.
- An appeal flow that re-runs classification with the submitter's counter-argument attached, under a second independent consensus round.
GenLayer Intelligent Contracts and the official toolchain:
genlayer-py,
genlayer-testing-suite,
genvm-linter and
genlayer-js. The frontend is Next.js 16,
React 19, Framer Motion and plain CSS design tokens.
MIT. See LICENSE.



