Case study · Sprout Solutions

The best onboarding team in the building was drowning in busywork

300+ hours per enterprise client, most of it work only a senior could do but nobody should have to. So we built agents for the boring 80% and left the judgment where it belonged.

300–350 hours per client2 prototypes3 onboardings shadowed
Role
Lead Product Designer
Company
Sprout Solutions · HRIS
Timeline
3 months, prototype + validation
Team
1 PM · 1 design · 2 eng

Here's an uncomfortable thing to discover about a team everyone loves: they're brilliant, clients rate them 9 out of 10, and they're being quietly buried alive by their own process.

Sprout's implementation team takes a freshly-signed client and turns “we bought payroll software” into “payroll actually ran correctly this month.” Consultants, translators and therapists rolled into one. On a 150-day enterprise onboarding, they'll spend 300 to 350 hours per client doing it.

High-skill work
15–20%of those hours are what the team is exceptional at
The rest
80%+parsing documents, validating spreadsheets, reconciling cells

Sprout's implementation team is its biggest onboarding strength. The process they operate within is its biggest constraint.

That sentence became the whole project. We weren't here to replace the team. We were here to give the busywork to something else.

01

What “implementation” actually means

When a company buys an HRIS, they don't get a product. They get a project.

Every client has its own payroll policies — overtime rules, 13th-month formulas, night-differential math, holiday stacking, government deductions — and almost none of it is standardised. An implementer has to extract those policies out of half-finished documents, translate them into configuration, run test payrolls, and explain — politely, repeatedly — why the client's beloved bi-weekly schedule has no tax table in Philippine law.

It's genuinely hard, judgment-heavy work. The problem was never the team. It was that the hard work sat buried under hours of work that wasn't hard at all — just tedious, manual, and unforgiving of a single typo.

The problem

Expertise trapped under admin

We shadowed the team through three onboardings. Three problems kept surfacing, and each made the others worse.

  1. High-skill work, zero assistance

    Reading a policy document and knowing instantly what will break lives entirely in the heads of three or four senior implementers. No tool, no checklist, no safety net. A junior walks into a kick-off meeting hoping they caught everything.

  2. Low-value admin eating the calendar

    Document review, spreadsheet validation and attendance reconciliation devour hours per client. One timekeeping variance run is about four hours of column-matching — and enterprise clients run it six times.

  3. No shared visibility across clients

    Each implementer juggles up to twenty concurrent clients with no dashboard and no early warning. By the time a client is “at risk,” it's usually already late.

The throughline: the team's expertise was real, but it was trapped — stuck in a few heads, and stuck under a pile of work that didn't need a human at all.

The bet

Agents do the 80%. Experts keep the judgment.

The temptation with AI is to aim it at the hard, glamorous part — “let the agent decide the policy.” That's exactly backwards, and it's how you lose a team's trust in a week.

Agents do the work that's tedious but mechanical. Humans keep every decision that needs judgment.

Agents parse, extract, compare and flag — then hand a human a clean, pre-digested starting point. The implementer still makes every call. They just don't start from a blank document and a cold spreadsheet anymore.

02

Catching the landmines before the meeting

Tool one · policy prep

Every payroll implementation starts with three documents, and today a senior reads all three by hand — two to three hours per client — looking for landmines.

BIR 2303
Tax registration certificate — legal entity, TIN, RDO code
IRD
Requirements doc from Sales — 40+ policy fields
SOW
Signed scope — modules, headcount, entities, go-live, frequency

A bi-weekly payroll frequency has no Philippine tax table — it goes back to Sales. A fiscal-year 13th monthisn't a blocker, but you'd better raise it at kick-off. Spotting those is exactly the expertise living in three or four heads, so we encoded it: the agent parses the documents, then runs every extracted policy through an escalation tree — return to Sales, flag for kick-off, or proceed.

Try it. Drop the three documents in, run the extraction, and watch the agent work through them.

Pre-Onboarding · Acme Corp

Document intake

BIR 2303·Tax Registration Certificate

Government-issued. Agent extracts: legal entity name, TIN, registered address, signatory, RDO code.

Drag & drop or click to upload

PDF

IRD·Implementation Requirements Document

Filled by Sales. Primary source for payroll policy matching. Agent checks completeness of 40+ policy fields.

Drag & drop or click to upload

PDF · XLSX · DOCX

SOW·Statement of Work / Design Proposal

Signed proposal. Agent extracts modules, headcount, entities, go-live target, payroll frequency.

Drag & drop or click to upload

PDF · DOCX

0 of 3 required documents uploaded

The point isn't the typewriter animation, fond of it as I am. It's the line at the end: 40 fields captured · 3 items need clarification. The implementer didn't read three PDFs. They got a triaged list of the three things that actually need a human — before they ever walked into the room.

The tool converts senior judgment into a team asset. That escalation-detection skill lives in three or four people today; encoded, every implementer catches what only seniors catch.

That's the whole thesis in one feature: expertise, un-trapped.

03

Reconciling two spreadsheets nobody wants to reconcile

Tool two · timekeeping variance

In the parallel run, the client runs payroll on Sprout alongside their old system. Attendance has to agree first, because every unresolved timekeeping variance becomes a payroll variance. Matching mismatched column names across 30+ attendance dimensions takes about four hours per company, per run — and an enterprise client with four sub-companies runs it six times. Roughly 96 hours of spreadsheet diffing on attendance alone.

Try it. Drop in both reports and run the analysis.

Step 1 of 3

Upload files

Drop both reports for this pay period. The Sprout file is computed from punches; the Client file is whatever they currently use to run payroll. We compare row-by-row.

Sprout — computed

Sprout Attendance Report

Generated by Sprout from raw punches for this pay period.

Choose Excel file

or drag and drop anywhere on this card

.XLSX · .XLS · max 10MB

Client — raw

Raw Attendance Report

Whatever the client currently uses — column names and layout can vary.

Choose Excel file

or drag and drop anywhere on this card

.XLSX · .XLS · max 10MB

0 of 2 reports uploaded

Notice what it doesn't do: it doesn't edit anything or pretend to resolve the variances. It hands back a colour-coded Excel with three sheets side by side and a blunt verdict — 11 to review, 6 blocking. Excel stays the editor, because that's where implementers already live. The agent just makes sure they walk into the variance meeting already knowing where the bodies are buried.

04

Why this wasn't an AI project

It's tempting to file this under “we added AI.” But the hard part was never the model. It was the domain.

Knowing that bi-weekly payroll is a hard blocker but a fiscal-year 13th month is just a flag. Knowing that a missing employee is worse than a two-peso rounding difference. Knowing that implementers want a triaged list, not an autonomous decision. None of that comes from a prompt — it comes from sitting with the team and learning the work.

The design system scaled the UI. Product knowledge scaled the judgment we encoded into the agents.

The agents are only as good as the escalation rules behind them, and those rules are senior implementers' instincts written down for the first time.

What success looked like

Defined before we built anything

This is a prototype in validation, not a shipped product with a quarter of funnel data — so no invented percentages. Instead, the bar we set and pressure-tested against:

Walk in prepared, every time

A junior catches the same pre-meeting landmines a senior catches — because the tool catches them, not their memory.

Hours back, not minutes

The 2–3 hours of document review and the ~4 hours per variance run are the target. Real numbers from real shadowing.

Agents assist, never decide

Every prototype hands over a clean starting point and stops. The moment it resolved a variance instead of surfacing it, we'd lose the team.

Quality must not regress

The 9/10 client scores are sacred. Saving time only counts if the work stays as good — or gets better because nobody's exhausted.

The real bar was never “automate the team.” It was give them their expertise back.

05

What I'd do differently

I'd put the ugliest possible documents in front of the intake tool on day one. Our prototypes shine on clean inputs — a tidy IRD, a well-formed export. The real world ships smudged scans from biometric devices, columns named OT_FINAL_v3_USE_THIS, and sub-companies whose policies contradict the parent. The agents earn their keep on those days. Next time I start with the worst file on the drive, not the demo file.

06

Key takeaways

  1. Point AI at the busywork, not the judgment. The fastest way to lose a team is to automate the part they're proud of.
  2. Encoding expertise is a product, not a side effect. The escalation rules are the feature.
  3. Domain knowledge is the moat. Anyone can call a model; almost no one knows that bi-weekly payroll has no tax table.
  4. “Hand back a clean starting point and stop” is a surprisingly powerful interaction model for agentic tools.

The best thing this project could do was make itself invisible.

So the team everyone already rated 9 out of 10 could spend all their time on the work that earned the 9, and none of it on the work that didn't.

This project was also the pilot for moving our design system off an npm package and onto copy-in components — act two of the Toge case study covers why that mattered.