How moderation works
Defense in depth across input, generation, catalog, alerts, and review.Flamify’s content moderation is built as defense in depth: multiple independent layers, each tuned to a different class of signal. Layers reinforce each other; no single layer is trusted alone.
The layers
1. Input — chat message moderation
Every user chat message is checked before it is sent to the AI worker. Stage 1 — keyword / regex matching Active forbidden topics are matched against the message using word-boundary keywords and regex patterns (typically 1–3 ms). Topics are admin-managed and versioned in production. Stage 2 — LLM verification When Stage 1 hits and LLM verification is enabled for that topic, a dedicated moderation model confirms whether the message is a genuine policy violation. The verifier is instructed to distinguish real violations from normal adult roleplay, BDSM language, and fictional framing. Per-topic confidence thresholds reduce false positives on noisy categories. Enforcement actions
Critical categories (minors / CSAM signals, self-harm, incest, zoophilia,
weapons instructions, drug manufacture, extortion, grooming/manipulation,
extreme gore) default to STOP. Softer categories (therapy substitution,
eating-disorder promotion signals, IRL meeting requests, AI-identity
confusion) default to HOLD.
The same pipeline runs on message edits. Regenerate reuses the prior
user text and does not introduce new unchecked input.
2. Character and catalog gates
- Characters cannot be created with age under 18 (schema-enforced,
ge=18). - Public catalog visibility requires an approved / published status.
- Characters carry an explicit
is_safeflag used by Safe Mode brand surfaces (see Safe filters). The flag is set by operators for app-store / PG-13 distribution — it is not inferred solely from generative output.
3. Generation-time controls
Chat → AI When a brand has Safe Mode enabled, AI requests are markedis_safemode=true
and routed with a Safe Mode chat configuration (model / parameter profile
intended for SFW surfaces).
Media generation
Under Safe Mode, the media pipeline:
- Strips NSFW tokens from free-text prompts (
nude,naked,nsfw, …). - Rejects explicit scene types.
- Prepends NSFW terms to the negative prompt (
nsfw, nudity, explicit, …).
child, underage, teen, …), independent of Safe Mode. The main
adult product intentionally allows adult nudity outside Safe Mode; age
blockers are never omitted.
4. Catalog and media distribution filters (Safe Mode)
Brand-aware Safe Mode (APP_PALM) filters discovery surfaces so that PG-13 /
app-store fronts only show characters marked is_safe and only mild media
tiers. Restricted media unlocks are blocked. Details:
Safe filters.
5. Alerts and human review
- Every STOP violation creates an alert log and notifies the trust channel.
- HOLD violations are logged; Slack noise is rate-limited and suppressed for expected adult-RP false-alarm categories where appropriate.
- Office / admin tools expose forbidden-topic CRUD and alert-log review so operators can tighten keywords, thresholds, and response templates without a code deploy.
- User and admin reports on media / characters trigger re-review against the current policy.
6. Community and operator reports
Users and operators can report problematic media, characters, or behaviour. Reports are prioritised by severity — suspected minor-related content first.Why this shape
Classifiers handle volume. Humans handle ambiguity. Safe Mode handles
distribution. Together they cover more ground than any single control.
What classifiers see vs what humans see
Automated classifiers process chat messages in-line because that is required to enforce Prohibited content. Classifier processing does not mean a human reads every private chat. Human reviewers see content that has been:- Flagged by a STOP / high-priority HOLD path,
- Reported by a user or operator, or
- Escalated as part of a trust-team investigation.
Fail behaviour
- If the moderation LLM is unreachable or rate-limited, the system may fail open on that stage (message allowed) to preserve availability — keyword Stage 1 still ran.
- Certain client/config errors on the verifier path fail closed (conservative block).
- Safe Mode media policy fails closed on explicit scenes (generation rejected).
What gets stopped most often
In rough order of severity priority (not necessarily volume):- Sexual content involving minors / grooming signals
- Self-harm and suicide ideation that triggers crisis routing
- Real-world weapons / explosives instructions
- Incest and zoophilia framings
- Extortion / blackmail language
- Drug manufacture / sourcing instructions