All work
Case 01 Verisk Germany · Oct 2022 - Present

0→1 agentic workflow, 85% faster claims

Founding designer for the agentic AI layer of Verisk Germany's claims platform. I ran UX workshops and user research, then set the information architecture, core workflows and design standards that govern it. It shipped on time under regulatory constraint and is bought today by some of Germany's largest insurers.

Representative interface: Tandem, a working agent console I designed and built end to end, where every agent states how far it may act on its own. Actual Verisk visuals are under NDA.

Role
Lead Product Designer
Company
Verisk Germany
Duration
Oct 2022 - Present
Impact
+85% productivity
Team
PM · PO · Developers · Design

Measurable gains, publicly reported

+85%Higher productivity through automation of previously manual tasks
hrs → minClaims processing time reduced from hours to minutes per case
−8 to −14%Sustainable reduction in time spent structuring audit reports for customers

All figures publicly advertised by Verisk. Resilience and scale verified in production. See verisk.com.

Give an agent a surface

The platform already worked. Claims came in, experts worked them, cases closed. Then AI services were added underneath it, and the interface had no way to hold them. The same system sits behind Medical Expert Platform.

So the brief was not "add AI to the platform". It was: give an agent a surface. Capability visible before it runs, scope adjustable by the person accountable for the outcome, and every result traceable back to the case it came from. (Final UI under NDA. See the Studio for comparable interaction work.)

A claim arrives carrying four documents, twenty-two pages, a claimant and an amount. An answer comes out the other side: one summary, one autonomy band, sent or held. Between them the agent reads, decides and acts, and that middle step is drawn as a wide empty box because there is nothing for the interface to render there. The output looks identical whether the agent read three documents or thirty. Under the empty box sit the four things the surface has to state instead: scope, confidence, sources read, and where it stops. What the interface had to render Nothing here to render A claim arrives Four documents, 22 pages, a claimant, an amount The agent reads, decides, acts An answer appears One summary, one band, sent or held So the surface states what it cannot show Scope Confidence Sources read Where it stops The output looks identical whether it read three documents or thirty
Why it's tricky
  • The capability is invisible. The output looks identical whether the agent read three documents or thirty. There is nothing to render, so the interface has to state the scope instead of showing the work.
  • Configuration is a permissions question wearing a settings UI. The person widening an agent's scope is a claims lead, not an admin, and the change carries legal weight. Every control has to carry its consequence with it.
  • The system is non-deterministic. Same input, different output. Standard UI patterns assume a repeatable state to design against.
  • "Relevant cases" is a claim, not a fact. When an agent cites precedent, the expert has to judge that relevance themselves, which means showing why a case was retrieved, not just that it was.
  • Trust has to be earned before the answer arrives. Once the result is on screen the reader has already decided whether to believe it, so the transparency work has to happen upstream of the output.

Validate quickly, ship on time

I ran UX workshops with the medical, legal and claims experts to surface what the work actually required. Then I audited the pipeline for where automated extraction broke down, and mapped user logic states against LLM context constraints with the engineers, so what we designed stayed inside what could be built.

Two user persona cards, every line of text blurred out: an avatar, a name and role, and three labelled fields on each card
Locked under NDA
The personas the workshops produced, held back. Two composite roles rather than one user: where each one's day actually goes, what they need the AI to do, and what breaks their trust in it. The permission split between them is the design problem, and it is the part I can describe. The sentences are under NDA, so the page carries a blurred capture and not the cards, and the words are not in the page source either.

Architecture flowcharts went to tested variants, then to production-ready components, with no detour through speculative mockups. The LLM constraints stayed visible at every step. (Specific flow diagrams under NDA.)

The delivery ran on an : product intent to structured design prompts, prompts to working prototypes, kept safe by design review.

Four steps: audit the pipeline, map user logic states against LLM context constraints, go straight to architecture flowcharts, then production-ready components. The speculative high-fidelity mockup stage is skipped. LLM context constraints stay visible across every step. 01 Audit the pipeline Where automated extraction was actually breaking down 02 · my step Map one layer deeper User logic states against LLM context constraints 03 Architecture flowcharts One shape design and engineering both read 04 Production components Built straight from the architecture, not a picture Speculative high-fidelity mockups skipped LLM context constraints stay visible at every step, not discovered at the end
The claim pipeline in six stages: intake, classify, extract, reconcile, decide, deliver, with three lanes running across them: what the assessor sees, what runs underneath, and what is read and written. The contents of each cell are withheld under NDA. Two of the six services are model-backed and marked in accent; the other four are ordinary deterministic software. One claim, six stages, and what runs underneath each one Cell contents under NDA Intake Classify Extract Reconcile Decide Deliver Interface Services Data Model-backed, and the only two stages that can be uncertain Ordinary software, deterministic, testable

The stages and the lanes are the shape of the work. What sits in each cell is the client's, so it is held back. The ratio is the part that mattered to the design anyway: the AI is two stages of six, and the interface has to say which two.

Customer acquisition outpaced the roadmap. I delivered with the PM against fixed dates, scoped each release to land on time without mortgaging the system underneath, and took the strategic calls with the team, through to consolidating the system end to end.

Choose speed or risk, deliberately

It is tempting to push for full automation, the fastest path, the most impressive on paper. But for migration, error checks, and high-liability cases, we could not risk AI hallucinations. We had to choose between speed and risk.

For deeper examples of how I approach these tradeoffs in unrestricted work, see the Studio.

NDA prevents naming the specifics. Here's how I categorize the judgment calls behind this work:

Craft that ships
  • Probed before designing. What is v1 actually for, what already exists that we can use, and what is the shortest path for this user group.
  • Knew what would ship and what would balloon. Front-end competence in the room meant "is this worth building?" came up at the sketch, not at code review.
  • Traded fancy components for legible and fast. Insurance experts want predictability, not delight. The UI had to disappear into the task. (Specifics under NDA.)
The Figma variables panel in Tandem: a Tokens collection of 27 semantic variables such as border, ring and destructive, each aliased to a primitive colour, with separate Light and Dark values
The three tiers, in the file. Tandem again, standing in for the NDA work: 31 primitives in one collection, 27 semantic tokens in another, and every token aliased to a primitive instead of a hex value. Light and dark are two modes of the same token, so a re-skin is a mode switch rather than a component change. Verisk's own libraries are under NDA.

Not the real UI. A complete redesign built for this portfolio, representing only services that are already advertised publicly.

Systems that scale
  • Fewer components, broader reach. Primitives that work across business units instead of a variant per team. Less to maintain, less to break.
  • Built for two years out. Domain knowledge told me which components would have to absorb new states, rules and integrations later.
  • Made room in the IA for business that did not exist yet. New products and AI surfaces could land without a re-platforming pass.
  • Tokens on a tiered architecture. Global, semantic, component. Engineering implemented once and we re-skinned across units without touching component code. (Taxonomy under NDA.)
One summary, three business units. Medical, motor and commercial liability are three different businesses with almost no field in common. Tab between them: it is the same build every time. Verisk's own units are under NDA.

Not the real UI. A complete redesign built for this portfolio, representing only services that are already advertised publicly.

AI with accountability
  • AI went where the roadmap needed it, not where the screens had room. Integration points came out of the PM plans, not out of retrofitting existing pages.
  • Regulated information, surfaced without slowing anyone down. Compliance carried by hierarchy and progressive disclosure, not by walls of mandatory text.
  • Refused to ship AI as a blackbox. What it was doing, how sure it was, and one click to the source. Trust was a design surface, not a copy decision. (Examples under NDA.)

Not the real UI. A complete redesign built for this portfolio, representing only services that are already advertised publicly.

  1. It classified the document, and said how sure it wasNot a filename and a spinner. The type it decided on and the confidence it decided with, before anyone asks.
  2. It concluded what it may do with the readingThe band is the agent’s own answer to that, not a setting a person left behind on the screen.
  3. It showed its work, per documentWhich agent read what, how many pages, how long it took, and how many fields it got out.
  4. It refused to pick between two documentsThe higher confidence is not the truer one, so nothing was written and the field says a person is needed.
  5. It declined to guess under the floorA read below 80% is left blank with its reason attached, rather than filled in and quietly believed.
A Tandem assessment after a run: four documents on the left, and on the right the agent's reading of them, one field held because two documents disagree and one left blank because it read below the confidence floor
The screen those decisions produced, with the machine's part of it marked. Every number is work a person would otherwise have done by hand and then had to be trusted about. None of it is a call the agent made alone.
Shipping with judgment
  • Pushed back on data-modeling shortcuts. Wrong choices in the data layer compound for years, and the interface pays the interest.
  • Only shipped features that moved efficiency. A competitor's cool feature had to name its underlying value first, or it got documented and shelved. (Specifics under NDA.)
  • Scoped engineering effort against design value, per feature. The team moved from "can we build it?" to "should we, and at what cost?" (Rubric under NDA.)

The assessment screen, built one capability at a time. Each version is the same page with features genuinely switched off, not a mockup of one.

Not the real UI. A complete redesign built for this portfolio, representing only services that are already advertised publicly.

Tandem: version one, one document and the verdict
V1 The form and the verdict
One document, and nothing else on the page. Nothing can contradict a single reading, so this version can never show a conflict. That is what V2 exists for.
Tandem: version two, the whole folder and its contradictions
V2 The folder, not the page
Added: the rest of the documents. Two of the four disagree, and the higher confidence is not the truer one, so the field is held and nothing is written.
Tandem: version three, attaching a document and re-running
V3 A folder that is not sealed
Added: attaching a document, and running it again. Attaching clears the old reading instead of silently re-running it, so a decision changes under the reader's own hands.
Tandem: version four, related cases in the rail
V4 What the person may see
Added: the related cases. A person may notice a pattern across three claims. The agent still reads only the one it was handed, which is why this came last.

Less interface, stronger system

The hardest interface problem in an agentic product is not what the AI does. It is what it declines to do. Every screen here that earned its place was built around a refusal: a field held because two documents disagree, a blank left where the read fell under the floor, a pattern across claims a person is allowed to notice and the agent is not. Confidence is easy to draw. Restraint is what people check you on.

And the front end kept getting smaller. Every time the model and the rules underneath it got stronger, a panel of controls stopped being necessary. Less to build, less to maintain, less for the assessor to read. Not every problem needs a new feature. Some need one fewer.

More case studies