Tax Assurance.
Designing VAT and SST returns a finance team can trust enough to sign. Built country by country across Saudi Arabia, Malaysia and France, then tested on a real prospect's own data.

A number someone signs
A VAT or SST return is a legal declaration with a person's name on it. Before filing, a tax team reconciles its books against the government's e-invoice records, finds what disagrees, and decides what to fix. Tax Assurance is the board they do that on.
Every design decision came back to one question: would a finance team sign a return based on this screen? I built it one country at a time: Saudi Arabia, then Malaysia, then France. I sat with each tax team, built the board with an agent, and validated it with them as I went. The agent wrote most of the code. I set the direction, made the calls, and judged the output. Then a real prospect tested the approach on their own data.
Every board shares one shape. It is the one thing I did not let vary by country.
The ledger on one side, the authority's e-invoice record on the other.
Matched, mismatched, missing, awaiting confirmation. Never two outcomes, never none.
So reconciliation and filing are one piece of work, not two tools.
From the number on the form to the documents behind it, filtered.
Almost everything on top of that chain changed by country.
| Malaysia | Saudi Arabia | France | |
|---|---|---|---|
| Return | SST-02 | VAT return | CA3, the 3310-CA3-SD |
| Reconciles against | MyInvois e-invoices | ZATCA records, with a Cleartax cross-check | Sales against the e-invoice and e-reporting feed. Purchases against vendor e-invoices |
| Rate logic | 6% and 8% by service type. The posted tax code decides the rate | One standard rate, 15%. The work is classification: the ERP tax code decides whether a line is standard, zero rated, export, exempt, GCC or a non claimable import | 20%, 10% and 5.5%, plus self-assessed import VAT |
| What the board had to add | Revisions after customer input, with a computed What changed screen | Eighteen coded findings drawn from how ZATCA rejects documents. B2C simplified invoices treated as reported, not cleared | E-invoice and e-reporting lanes, credit notes against prior periods, payment confirmations, a per document audit trail |
| What needs a human | Unrated recharges and rate variances, until confirmed | Each finding arrives with its own what to do. Data quality exceptions sit apart from reconciliation gaps | Outcomes marked review or resolve. Payments awaiting confirmation |
| Stays identical | Navigation shape, the drill pattern, uncertain value kept out of totals, status as a word, one table treatment | ||
Built on CTAI
CTAI is Clear's AI platform. Finance teams build agents on it to solve their own problems. Every Tax Assurance board lives there.
Boards are not shipped by an engineering team. A builder works with an agent: the agent writes the queries over the account's data and the board itself, a single file app, and the platform publishes it to a shared library. A tax team opens it there, shares it, and can question it in chat.
That changes where design sits. There is no design file and no handoff. Whatever the builder and the agent produce is what the customer sees. So the design system is not a reference document. It is the only thing between a builder and a board that looks like nobody designed it.
The platform is Clear's. The boards, the patterns and the system that renders them are mine.
Saudi Arabia: where the patterns came from
The Saudi board builds the VAT return two ways, from the ledger and from the e-invoices cleared or reported to ZATCA, and shows where they disagree.
My manager pointed out that the first version had three problem surfaces, each drilling somewhere different, and none telling the user what to do. So every figure now drills to its documents, each with a what to do line, and the board opens on a verdict: not ready to file, how many documents are holding it up, and how much VAT is at risk.
Build it in, or hand it to the agent
Auditing the e-invoices against ZATCA's own cleared feed is a real need, but that feed is not held in the workspace. Building it into the board would mean ingesting data the board does not own, for a check that runs during an audit, not every month.
So the board explains when the check is needed and writes the prompt for it. The user exports the ZATCA report, opens a chat, attaches it and pastes the prompt. The agent matches both sides and builds its own report. The boundary is stated on the screen: the agent's work does not change this filing.
The same board has a small ask box for quick questions, with the board's computed figures passed to the model as context.
My part was deciding what the board hands to an agent, and what it keeps away from one.
Give the agent the one-off work. Keep it away from the number being filed.
Malaysia: two dashboards became one product
That combined board became Tax Assurance, and it is why every board has the shape shown at the top of this page.
Rules about numbers
Malaysia is where the most important decisions turned out to be about what a number means.
0.00 means two sides agree. A dash means nothing here. n/a means no counterpart can exist.
The service mapping only groups, so it can never move a filed figure. A conflict raises a finding instead.
Shown in the books' sign convention on both sides. Otherwise a RM 677 gap reads as RM 18,000.
The numbers are proved, not trusted. A separate classifier re-derives every document's status from the raw rows, and a test asserts it matches the scenario that generated it, for all 1,255 documents. It caught 27 real bugs. One: adding 25 days to the 2nd of a month stays in the same month, so a document cleared next period was silently marked matched.
Then the customer replied with two kinds of correction. Three mapping corrections moved the return by RM 0.00, because the tax code sets the rate. Confirmed re-codings in the ledger moved it by RM 6,522.10.
Stating the change, or computing it
The quick way to show a revision is to overwrite the numbers and add a note. That is how a return quietly drifts from what was actually filed.
So revision 1 stays frozen as filed, and revision 2 is a separate board. The build fails if any revision 1 file changes. A What changed screen computes the difference from both datasets and puts both kinds of correction side by side, so the customer can see why some of their changes moved the return and some could not.
A filed number never changes quietly.
France: the return as a declaration
France is moving to mandatory e-invoicing, so a single sale can go wrong in more ways than a missing invoice. It can be sent down the wrong lane, rejected by the buyer, or reported without its payment ever being confirmed. Every finding rolls up into one of seven outcomes, each marked review or resolve.
The form shows the books. The findings live elsewhere.
The tempting design was to put every difference on the return, beside the box it affects. That turns a statutory form into a worksheet, and the person signing it can no longer tell what they are declaring from what is still disputed.
So the CA3 screen shows the return as the books state it. Differences live on the recon summary. Boxes backed by ledger lines drill into them, and only those boxes carry a link. Every screen ends on the same line: prepared draft, nothing filed or transmitted.
A return should read as a declaration, not a worksheet.
The live board needs a wider screen. Open it on a laptop to click through all nine screens.
Then a real prospect
Everything above runs on synthetic companies. The real test came from a sales call with the finance head of a Malaysian device distributor who used Clear only for e-invoicing. Every sales question took half an hour of exports and pivot tables. SST was filed by hand. And any new tool meant months of IT and security review.
The agreed proof of concept: show something real, from data already in Clear, with no IT involvement. Gaps were acceptable, as long as they were shown rather than hidden. That meant building the SST return from e-invoices alone, with no ledger at all.
The findings are sorted by what they do to the return. Credit notes that cannot support their deduction, tax charged on only part of a line, a service tax rate to confirm, imported services with no tax. A fifth group, e-invoice compliance issues, is labelled as having no effect on the return, so it never gets mistaken for one.
Two of my early claims were wrong: an imported services figure left in unconverted foreign currency, and a possible under-declaration that turned out to be supplier credit notes. The checks caught both. The corrections went onto the screen and to the stakeholder, with the corrected figure and the cause. With a prospect deciding whether to trust you, the correction is part of the product.
The same e-invoices then answered the question the finance head asked on the call: how many units of one phone model sold between two dates. That had been a half hour of exports. It became one screen, inside a fourteen screen sales view built from nothing but e-invoices.
The largest customer appeared under two names. A pivot by name would have split it in two.
No payment data, so the cash calendar says what falls due, not what is unpaid.
Almost all of the pricing leakage was one customer's contract pricing. That customer is now shown apart, with the reason.
The overview loads its own data first and the rest in the background. First paint went from 6.5 to 3.0 seconds, measured.
Making it repeatable
By the time France was done, I had written three skills that each claimed to be the design system. All three were mine, written at different points along the way. They disagreed on more than twenty concrete points, and one contradicted itself, still describing a teal brand after the tokens had moved to near black.
That was survivable while I was the only builder, because the real rules lived in my head. It would not survive PMs and engineers building with it, or an agent following it literally.
The tokens were identical across boards. What had drifted was the documentation, the thing an agent or another builder reads to reproduce the system.
Rule once, or let every builder decide
The easy fix was to merge the three skills and soften the conflicts into guidance. That works only while one person knows which rule is the real one.
So I ruled. Fourteen conflicts, each decided once, each with a reason, and the system cut from 72,000 characters to a manifest of about 160 lines.
Filters are the clearest example. Boards used to carry a filter bar on top, and whoever built the board chose what went in it. It cost space, and users pushed back. These tables have many columns, and they wanted to filter on whichever one mattered to them. So filtering moved into the column headers. Every column filters, and the bar is gone.
A design system that does not rule is a suggestion.
| Conflict | Ruling | Why |
|---|---|---|
| Filters | Column headers filter. No filter bar above the table | Users wanted to filter on any column, not the few a builder picked. The bar also cost space |
| Status | A chip with a word. A dot only alongside text | Meaning must never ride on colour alone |
| Uppercase | Forbidden | A dashboard is read, not shouted |
| Tables | Two variants only, one per board | Mixing them makes a board read as two products |
| Colour named tokens | Retired | A colour name carries no meaning and breaks on repalette |
| Generic design advice | This skill wins for dashboards | Otherwise the generic skill keeps firing on the word "dashboard" |
I had built Malaysia by copying the Saudi board, and every copy carried the last country's assumptions. So the two concerns were split. A logic agent holds everything that belongs to a country: it maps an account's own tables onto a fixed column contract and a common tax code vocabulary, and checks the maths. One rendering skill draws everything a person sees. The acceptance test: the Saudi board, rebuilt through the skill alone, had to reproduce the hand built board byte for byte.
In a single file dashboard, almost every design mistake is silent. One selected control rendered white on near white inside the platform while correct locally, because the platform's injected theme overrode ours. No review could have caught it. So conformance became code: eleven static checks, and fifteen more in a real browser against a deliberately hostile theme. Every check is itself tested: break what it guards, confirm it fails, then revert. "It looks right" is not evidence.
$ python audit_render.py france_vat_assurance[PASS] the page declares a doctype standards mode[PASS] the board's own values survive the host theme identical under both[PASS] every screen renders with no JavaScript error 8 screens[PASS] the content pane scrolls scrolled to 381px[PASS] the nav stays pinned while the content scrolls fixed[PASS] every selected control meets WCAG AA (4.5:1) .segbtn.on = 17.76:1[PASS] every descendant styled class has its ancestor 11 classes across 8 screens[PASS] every table row matches its header 8 screens walked[PASS] charts render at non zero size, leave no orphan canvases 2/2, live 2[PASS] the grid is visually identical to .table 8 computed properties[PASS] every export produces a populated sheet 5 sheets[PASS] the page never scrolls horizontally 1366, 1180, 1024, 820
Some things no check can catch: a summary that repeats its own detail and drifts from it, a row with two things to click, a dense screen that reads as clutter. Those came back from review, and they always will. The audit is a floor, not a substitute for looking at the page.
Where it landed
Tax Assurance boards for Saudi Arabia, Malaysia and France, in market with six customers across the three countries. A working proof of concept on a real prospect's data. And one design system with runnable audits, the standard for every Tax Assurance board and now used by the PMs and engineers who build with it.
Next is a scheduled agent that emails a monthly management pack and alerts from the same tables. It is planned, not built.
What is still open
Not every board follows its own rules yet: the French board has no verdict or agent handoff, and two headline cards show at-risk figures without naming their scope. And the audits measure what a board has, never what it used to have, so a dropped feature is still caught by a person.
What I took from this
Decide a finding once, where the data lives.
Then no two screens can disagree about it.
Account for all of it before using any of it.
An incomplete answer that shows its gaps is more trustworthy than a complete one that hides them.
Define what must be true, then let it be built.
When agents and other builders do the building, design leadership moves from reviewing screens to defining what must be true, and knowing which parts only a person can judge.
What it adds up to.
Every company shown in this case study is invented and every figure generated, except the prospect section, where all visuals are masked and the business is described generically.
Back to all work