Start with real public conversations, switch to the written training set, or paste your own. The rules engine scores the agent turns against a 7-part scorecard and 5 auto-fails, checks stated facts against the help center, and, when the customer links a market, against that market's live rules, fee schedule, bond and challenge window. The AI reviewer gives a second opinion you can compare.
Conversation
Findings and who fixes them
AI reviewer (second opinion)
@AskPolymarket is Polymarket's public AI bot: it answers posts on X with the current odds and a market link. This tab takes its recent replies, resolves each link to the event, matches every quoted outcome to a market, and compares the quoted percentage with the CLOB price at the minute the reply was posted. A real accuracy audit of a real AI agent, from public data.
Verified: within 2 points of the price at that minute (the bot rounds to whole percent). Wrong: the figure maps to exactly one market and is more than 2 points off. Ambiguous: several markets fit the wording (for example quarter-final and semi-final markets for the same team), so the figure is not scored against the bot. Price history: clob.polymarket.com/prices-history at 1-minute fidelity.
Everything the rules engine does, in plain words, with the source each rule comes from. A QA program people trust is one whose rules they can read.
Auto-fails
Fact checks
Rules engine vs the lead's reference scores (training set)
Data sources and limits
- Live X conversations: api.fxtwitter.com (public posts only, fetched in your browser). Polymarket's private chat and DM conversations are not public, so live scoring uses public replies.
- Help center: help.polymarket.com sitemap and articles. Docs: docs.polymarket.com Markdown pages. Market data: gamma-api.polymarket.com. Prices: clob.polymarket.com. Deposit minimums: bridge.polymarket.com.
- The training set is written for the demo and labelled as such. The rules engine is deterministic pattern matching; the AI reviewer (Llama 3.3 70B on Workers AI) is a second opinion and is shown side by side, never merged silently.
- Nothing you paste is stored. The worker keeps a 10-minute to 6-hour cache of public pages only.
Monthly calibration: four graders score the same conversation, and the rules engine and AI reviewer score it too. Criteria where scores spread by 2 points (on a 0-2 scale) go on the calibration agenda. Edit any score to run the session live.
Calibration agenda
AI-assisted QA reviews every conversation instead of a 2-3% sample. Below, the rules engine has scored all sample conversations; the calculator shows what that coverage means at real volume.
Coverage calculator
By channel
Failure patterns
The same knowledge feeds human agents and the AI agent, so a wrong article is wrong everywhere at once. These checks read the help center and docs live and compare them with each other and with the live market data from the Gamma API.
Help-center review queue
Articles last edited before the 28 Apr 2026 exchange upgrade (new contracts, pUSD collateral, fees in USDC at match time) that mention trading, fees, funds or resolution.
Every finding gets one owner. Agent behaviour goes to coaching, repeated patterns become a training module, knowledge defects become a content ticket, and anything the AI agent got wrong becomes a knowledge or guardrail patch. Run the Knowledge health checks first to include content defects.