CX Quality Loop

QA, knowledge and training as one system for prediction-market support: score human and AI conversations, calibrate graders across in-house and BPO teams, check the help center against live market data, and turn every failure into a coaching note, a training module or a content fix.

Independent prototype built by Edward Tay for a job application. Not affiliated with or operated by Polymarket. Live conversations are real public replies from @PolymarketHelp and @AskPolymarket on X; the training set is written for the demo and labelled. Help-center text, docs, market data and prices are fetched live from Polymarket's public pages and APIs.

Start with real public conversations, switch to the written training set, or paste your own. The rules engine scores the agent turns against a 7-part scorecard and 5 auto-fails, checks stated facts against the help center, and, when the customer links a market, against that market's live rules, fee schedule, bond and challenge window. The AI reviewer gives a second opinion you can compare.

Conversation

Rules engine
-
Lead's score
-
calibrated reference

Findings and who fixes them

    AI reviewer (second opinion)

    Llama 3.3 70B on Cloudflare Workers AI grades the same scorecard. Use it to widen coverage; the lead's calibration decides when it can be trusted on each criterion.

    @AskPolymarket is Polymarket's public AI bot: it answers posts on X with the current odds and a market link. This tab takes its recent replies, resolves each link to the event, matches every quoted outcome to a market, and compares the quoted percentage with the CLOB price at the minute the reply was posted. A real accuracy audit of a real AI agent, from public data.

    Verified: within 2 points of the price at that minute (the bot rounds to whole percent). Wrong: the figure maps to exactly one market and is more than 2 points off. Ambiguous: several markets fit the wording (for example quarter-final and semi-final markets for the same team), so the figure is not scored against the bot. Price history: clob.polymarket.com/prices-history at 1-minute fidelity.

    Everything the rules engine does, in plain words, with the source each rule comes from. A QA program people trust is one whose rules they can read.

    Auto-fails

    Fact checks

    Rules engine vs the lead's reference scores (training set)

    Data sources and limits

    • Live X conversations: api.fxtwitter.com (public posts only, fetched in your browser). Polymarket's private chat and DM conversations are not public, so live scoring uses public replies.
    • Help center: help.polymarket.com sitemap and articles. Docs: docs.polymarket.com Markdown pages. Market data: gamma-api.polymarket.com. Prices: clob.polymarket.com. Deposit minimums: bridge.polymarket.com.
    • The training set is written for the demo and labelled as such. The rules engine is deterministic pattern matching; the AI reviewer (Llama 3.3 70B on Workers AI) is a second opinion and is shown side by side, never merged silently.
    • Nothing you paste is stored. The worker keeps a 10-minute to 6-hour cache of public pages only.

    Monthly calibration: four graders score the same conversation, and the rules engine and AI reviewer score it too. Criteria where scores spread by 2 points (on a 0-2 scale) go on the calibration agenda. Edit any score to run the session live.

    Calibration agenda

    AI-assisted QA reviews every conversation instead of a 2-3% sample. Below, the rules engine has scored all sample conversations; the calculator shows what that coverage means at real volume.

    Coverage calculator

    By channel

    Failure patterns

    The same knowledge feeds human agents and the AI agent, so a wrong article is wrong everywhere at once. These checks read the help center and docs live and compare them with each other and with the live market data from the Gamma API.

    Help-center review queue

    Articles last edited before the 28 Apr 2026 exchange upgrade (new contracts, pUSD collateral, fees in USDC at match time) that mention trading, fees, funds or resolution.

    Every finding gets one owner. Agent behaviour goes to coaching, repeated patterns become a training module, knowledge defects become a content ticket, and anything the AI agent got wrong becomes a knowledge or guardrail patch. Run the Knowledge health checks first to include content defects.

    QA finding→root cause→coaching / training / content / AI patch→re-audit next week

    Training module draft

    Weekly quality report