Review methodology
This page documents how I actually test the AI companion platforms I cover on janitor-ai.io. It's not a marketing document — it's the working checklist I use before publishing a review. If a review on this site doesn't follow this methodology, it either isn't a full review yet or the deviation is noted at the top of the piece.
Testing time
Minimum two weeks of active use before any review gets published. In practice this usually stretches to three or four. Two weeks is the floor because that's how long it takes for a platform's honeymoon quality to fade — the initial "wow" of a first character interaction rarely holds up past day five. The interesting behaviour shows up around day ten, when memory drift, model quirks, and moderation edges start surfacing in real conversation.
For Janitor AI specifically, that means at least a fortnight of daily sessions across multiple characters, with both JanitorLLM and at least two BYOK models (typically DeepSeek and Claude Haiku, sometimes GPT-4o for comparison).
Model coverage
Every Janitor AI review tests four model configurations at minimum:
- JanitorLLM (free tier). The default. Tests the actual out-of-the-box experience for someone who just signed up.
- DeepSeek via BYOK. The cheapest premium path. Community consensus favourite.
- Claude Haiku via BYOK. Balanced quality-to-cost, popular in the r/JanitorAI budget threads.
- GPT-4o via BYOK. The high-end reference point.
Each configuration runs the same test scenarios so the model differences show up clearly. The scenarios cover casual conversation, slow-burn roleplay, high-stakes scene handling, and long-context memory retention over 100+ turns.
What we measure
Six criteria, evaluated separately, weighted according to how much they actually matter in real use:
- Writing quality. Coherence, tone matching, prose rhythm, personality consistency. Judged by reading, not by scoring rubrics.
- Memory reliability. Does the character remember specifics from earlier in the session? How far back?
- Uptime and reliability. Measured across the two-week window, with specific attention to the 7 PM – 1 AM UTC peak-hour window.
- Safety and moderation. Where filters fire, how they're implemented, and whether the platform is honest about them.
- Setup friction. How long from landing to first character reply. How complicated is BYOK configuration.
- True cost. Free tier ceiling plus realistic BYOK monthly spend, benchmarked against Reddit budget threads.
Cross-verification
No claim in a review gets published based only on my own testing. Community reports on r/JanitorAI and r/JanitorAI_Official, Discord announcement archives, and — for pricing — actual API provider dashboards get cross-referenced before anything ships. If a claim can't be verified against at least one independent source, it either gets flagged as opinion or dropped.
Update cadence
Every review gets re-verified at a minimum of every 90 days. Product features move fast in this category — a review that was accurate in April is often factually wrong by July. On re-verification, the dateModified stamp updates, the "Last reviewed" date at the bottom of the page updates, and any material changes get called out in a changelog block at the end of the review.
Transparency notes
Two things worth naming explicitly. First: I'm one person, not a testing lab. There's no A/B statistical rigor here — the methodology is designed to catch obvious issues and stress-test the marketing claims, not to produce peer-reviewable science. Second: my own preferences bias what I notice. I write fiction as a hobby, so I probably weight prose quality higher than a user who just wants small talk would. That bias is worth knowing about.
The affiliate disclosure that governs how any of this connects to revenue is on the editorial policy page.
Where this methodology applies
The current review that runs on this methodology is the Janitor AI coverage on the homepage — that's the site's flagship editorial output. Context on who runs janitor-ai.io and its third-party posture toward the platform is on the about page. Author details are on the editor's profile. Site-wide legal boundaries are in the terms, privacy policy, age verification, and the § 2257 statement.
Last reviewed: July 9, 2026 · by Melissa Blake