Start here
This bluebook trains one capability: receive an unfamiliar CS Ops task packet, rebuild its logic from the evidence, decide whether it reflects real work, check the analysis yourself, and explain exactly what is wrong.
A reviewer decides whether a CS Ops task is realistic, solvable, operationally correct, and supported by its data and answer key.
Where each part is taught
Tap any part to jump to its chapter.
How the skills stack
How to use this bluebook
Chapters 1 to 3 build the reviewer's operating system. Chapters 4 to 8 give you just enough CS Ops depth to judge a task, each ending with a reviewer lens. Chapters 9 and 10 are answer-key forensics and a full practice packet. Chapter 11 prepares your application answers.
Progress is saved in this browser. Finish every chapter to unlock the final test and certificate.
The measure of success
"I don't need to know this company's systems beforehand. I can rebuild the operating logic from the evidence, judge whether the task resembles real CS Ops work, check the analysis myself, find what the AI got wrong, and say precisely why."
The reviewer's job
You receive a task packet: a short request in a colleague's voice, a set of exports, and an answer key. Your job is to decide whether it holds up, and to say why in one line.
In simple words
You are the examiner who checks an exam paper before students sit it. Is the question clear? Is the information enough? Is the answer key actually right? Would a real person in this job ever be asked this?
The eight questions behind every review
Who owns what in real CS teams
Use this to judge "would a CS Ops analyst be asked this?". Realistic tasks ask for analysis, not negotiation, selling or legal decisions.
| Role | Owns | Typically asks CS Ops for |
|---|
What real CS Ops work looks like
Words you will meet in every packet
| Term | Meaning |
|---|
The reviewer reasoning engine
Every packet goes through the same eleven steps. The order matters: you define the unit, population and timeframe before touching numbers, and you solve the task yourself before reading the key.
In simple words
A good judge doesn't read the verdict first and then look for evidence. They read the case, check the evidence, reach their own view, and only then compare it with what was decided.
Walk the eleven steps
Pick a step to see what to ask, the usual trap, and how it applies to the practice packet in chapter 10.
Grain first, always
Before any join or total, say what one row means in each file: one contract, one account per week, one ticket, one invoice line. Most wrong answer keys come from mixing grains.
Confidence is part of the answer
State it as High, Medium or Low. Low confidence is honest when the data is thin; it tells the researchers which packets need another look.
Task realism and data realism
A beautifully written task can still be unsolvable because a required field is missing, or unrealistic because no real system would produce that export.
Is the request realistic?
Weak request
"Find customers who might churn."
No population, no timeframe, no rule, no output. Many answers are defensible.
Realistic request
"Identify contracts ending within 120 days whose non-renewal notice deadline has not passed, and flag accounts showing at least two independent risk signals."
Population, window, rule and output are all explicit.
Inspect each export
Pick a file type to see its grain, key fields, joins and the traps a reviewer looks for.
The ten checks for every file
| Question | What the reviewer checks |
|---|
Which source is the truth?
When files disagree, the answer key should rely on the source that owns the fact.
| Fact | Source of truth |
|---|
Reading NPS and support data correctly
The renewal queue
The renewal queue is the list of contracts that need action now. It is built from the contract register, not CRM close dates, and the notice period decides the real deadline.
In simple words
A phone plan that renews by itself only lets you cancel if you tell the company 30 days before. Day 29 is too late. The queue makes sure the team acts before that last day, starting earlier for the biggest customers.
The fields that decide the window
| Field | Why it matters |
|---|
Run the queue on sample contracts
Change the "as of" date and watch each contract move. Lead time is 120 days for $100,000+ ARR, 90 days for $25,000 to $99,999, and 60 days below that.
Renewal queue builder
| Account | ARR | End date | Auto-renew | Notice | Action deadline | Queue opens | Status |
|---|
Action deadline = end date minus notice period for auto-renew contracts, and the end date itself for manual renewals. Queue opens = action deadline minus lead time.
Edge cases that break naive queues
Tap a card to see how an expert handles it.
Uplift, proration and true-ups
Renewal prices come from the contract: the uplift, the cap, the committed quantity and the true-up clause. An answer key that ignores any of them is wrong.
In simple words
Uplift is the yearly price rise, like rent going up, but the contract may say "never more than 5%". A true-up is like a buffet charging for extra guests: you agreed to 100 seats, 118 people used it, so you pay for the 18 extra.
Check a renewal calculation
Renewal and true-up calculator
The proposed 9% uplift is above the 7% cap, so the cap applies. Renewal ARR assumes the new commitment matches actual usage.
Commercial adjustments: ticket size, discounts, waivers and extensions
Real renewals are rarely list price times uplift. These adjustments appear in packets often, and answer keys often get them wrong.
| Adjustment | Correct treatment | Typical answer-key mistake |
|---|
Contracts differ. When the files don't say how an adjustment is treated, that is a definition gap to flag, not a guess to make.
Health signals and escalation triggers
A trigger says "a human must look at this account now". A reviewer checks that the rule is realistic, corroborated, and tested for both false alarms and missed risks.
In simple words
A smoke alarm is useful because it rings for smoke. If it rings for toast, people stop listening. If it stays silent during a real fire, it is worse than useless. Good triggers avoid both.
From signal to escalation
Tune a health score
This sample account scores Usage 40, Adoption 55, Support 70, NPS 30 and Billing 100. Change the weights and watch the verdict move: a reviewer asks whether the weights were tested against real churn.
Health score builder
55 / 100Biggest drag on the score: Usage
The trigger library
| Trigger | Source | Typical rule | Severity |
|---|
Thresholds are starting points. Every company tunes them to its own churn history.
Test an account against the triggers
Tick what you observe. The tool scores the risk and lists what to rule out first.
Escalation checker
GreenMonitor: no escalation yet. Keep the account on the weekly watch list.
False positives: the alarm rings for toast
Tap a card to see why the trigger may be wrong and how to check.
False negatives: the fire nobody saw
Accounts the triggers miss. Tap a card to see why and how to catch it.
Churn, retention and reconciliation
The CRM, the CS platform and billing each give a different churn number until someone reconciles them. Packets test whether the answer key knows which number is right.
In simple words
Your notebook says you have 500. Your wallet has 450. Reconciliation is finding the missing 50: the snack you forgot to write down, or the coin you counted twice. Every difference has a reason.
Gross and net revenue retention
Retention calculator
Same customers, start to end of periodNew customers signed during the period never count in either. GRR can never exceed 100%.
Why CRM and billing disagree
Pick the mismatch you see to get the usual cause and the fix.
The reconciliation method
CS platforms and billing systems
You don't need to administer Gainsight or Zuora to review a task. You need to know enough about how they work to tell whether an export or request is plausible.
In simple words
Different car brands put the buttons in different places, but every car has a steering wheel, brakes and a fuel gauge. The CRM is the promise, billing is the receipt, and the CS platform is the dashboard that reads both.
CS platforms: same jobs, different names
| Job | Gainsight | ChurnZero | Totango | Planhat |
|---|
Vendors rename features often. Check current documentation before disputing a name.
Billing systems: the objects you will read
| Concept | Zuora | Chargebee | NetSuite | Stripe |
|---|
Billing events and what they mean for CS
| Event in billing | What it means | Correct CS Ops treatment |
|---|
Answer-key forensics
"Customer ABC is at risk because usage fell 42%." A reviewer doesn't only ask whether 42% is right. They ask whether the conclusion is supported at all.
In simple words
A detective doesn't accept a confession just because it sounds right. They check whether the evidence supports it, whether someone else could have done it, and whether the timeline fits.
Ten questions for every answer line
The nine ways a task goes wrong
Tap a failure type for an example and the fastest way to spot it.
Your verdict options
| Verdict | Use it when | Example one-line reason |
|---|
Writing the one-line reason
Weak
"The answer is wrong; Evergreen shouldn't be there."
Expert
"Dispute: Evergreen Schools' 30-day notice deadline (1 Oct) passed before the as-of date (3 Oct), so it auto-renews and should not be in the actionable queue."
Spot the flaw
Each card describes something found in a packet. Tap for the verdict and the reason.
The reviewer scorecard and a practice packet
Don't give a packet one overall score. Judge each dimension separately, then review a full packet end to end.
The reviewer scorecard
Set each dimension. The tool drafts your form summary.
Packet scorecard
Accept| Dimension | What it asks | Your call |
|---|
Practice packet: the Q4 renewal risk list
Read the request and files, solve it yourself, then compare with the AI answer key and choose your verdict.
The request
From Priya, VP Customer Success: "As of Saturday 3 October 2026, which renewals do we still need to act on before 31 December, and how much ARR is at risk? Use our standard rule: an account is at risk if it shows two or more signals."
Signals: weekly active users down 30% or more against the 90-day baseline; under 50% of paid seats active; two or more P1 tickets in 30 days; an NPS of 0 to 6 from a decision maker; an invoice 30 or more days overdue.
AI answer key
At risk: Northwind Labs $84,000; Bluefin Media $30,000; Cedar Health $135,000. Total ARR at risk: $249,000. Dover Freight excluded because its contract ends in February.
Your verdict
Choose a verdict, then reveal the expert review.
Writing new tasks from your own experience
| Part | What good looks like |
|---|
Answering the expert questions
Application questions test whether you have done the work. Answer with exact fields, thresholds and failure cases.
Use these honestly
These model answers show the structure and depth reviewers look for. Rewrite them with the systems, numbers and stories from your own career. Never claim a tool or project you have not worked with: the domain test and the call will probe the details.
How did you build the renewal queue at your last company? What fields decided the window?
Describe your escalation triggers for at-risk accounts, and cases where they can be wrong.
How did you reconcile churn and expansion in the CRM to billing?
The answer formula
The 14-day prep plan
Build a small dataset, break it on purpose, and review packets against the clock until spotting errors feels automatic.
Your progress
0 of 15 tasks done. Progress is saved in this browser only.
The daily 60-minute practice
The best reviewer doesn't know every company's systems. They can rebuild the logic from the evidence, and explain in one line exactly where a task breaks.
Final test and certificate
Twenty-two practical questions, mostly reviewer judgement. Score 80% or more to earn your certificate.