Turning risk alert reviews into an ML feedback loop

How I moved frontline compliance judgment into the alert-review moment while Behavox analysts retained control of model tuning.

Compliance officers had the clearest context to judge whether an alert was accurate, but their judgment sat outside the tuning path. I designed an in-context feedback path that captured their assessment beside the alert evidence and sent it to Behavox analysts.

My role and scope

My role and scope

Product Designer: Working with product a manager, a researcher, and two engineers

Product Designer: Working with product a manager, a researcher, and two engineers

Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery

Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery

Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM

Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM

64%

Boost in risk alert accuracy

Boost in risk alert accuracy

31%

Increase in usability score

Increase in usability score

Improved ML training loop

Improved ML training loop

SOLUTION

Compliance officers gained a direct feedback path

Officers gained a direct feedback path

The launched workflow kept feedback inside alert review. It connected frontline judgment to analyst tuning input without transferring control over model changes.

Validate detected signals

Validate detected signals

A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.

1

1

Evidence stays in view

Evidence stays in view

Evidence stays in view

2

2

Feedback sits beside the decision context

Feedback sits beside the decision context

Feedback sits beside the decision context

Classify additional signals

Classify additional signals

End-to-end workflow showing highlighted text linked to a compliance scenario via an inline menu and a side panel form.
Inline context menu on highlighted email text with options to create a label or link the content to a compliance scenario.
Side panel form for linking highlighted content to a compliance scenario, with dropdowns for scenario and use case and save actions.
Side panel form for linking highlighted content to a compliance scenario, with dropdowns for scenario and use case and save actions.

1

Classification starts from selected evidence

Classification starts from selected evidence

2

2

New signals link to existing scenarios

New signals link to existing scenarios

1

Classification starts from selected evidence

2

New signals link to existing scenarios

40

Manager oversight stayed separate from the launch

The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.

The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.

A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A UI component labeled "Risk alert justifications" displaying a flagged message: "I can send you a $500 thank-you if you approve the renewal before Friday," identified as an indicator of an "Improper Inducement – Offer to Bribe" risk, with buttons to validate the signal as accurate or inaccurate.
An expanded UI component showing "Risk alert justifications," detailing the signal reason, risk type, policy reference, and specific metadata like matched phrase and participant count for a flagged "Improper Inducement – Offer to Bribe" risk.

The documented concept covered alert accuracy, validation coverage, review activity, and trend-level monitoring for later delivery.

SOLUTION

The ML feedback loop now starts inside the review flow

I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.

SOLUTION

The ML feedback loop now starts inside the review flow

I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.

BEHIND THE SOLUTION

The simple review action depended on a harder sequencing decision

BEHIND THE SOLUTION

The simple review action depended on a harder sequencing decision

BEHIND THE SOLUTION

The visible change was small. The sequencing decision was the real work.

The harder question was what to ship first across three roles with different responsibilities and data needs.

Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.

Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.

The harder question was what to ship first across three roles with different responsibilities and data needs.

PROBLEM

Officer judgment lived outside the tuning path

PROBLEM

Officer judgment lived outside the tuning path

PROBLEM

Officer judgment lived outside the tuning path

Feedback required a separate step after review, delaying when analysts could use the signal.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise, but the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it. Managers also lacked a reliable aggregate view of review coverage and model health.

Feedback required a separate step after review, delaying when analysts could use the signal.

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.

Legacy tuning path:

Legacy tuning path:

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.

Legacy tuning path:

Legacy tuning path:

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise. Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.

Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."
Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

EXPLORATION

I evaluated new ways to capture officer judgment

Three key concepts varied placement and response depth before testing moved feedback into risk alert justifications.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Structured popover: guided questions attached directly to selected evidence.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: strong context, but feedback appeared before officers' moment.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Guided signal review: structured input inside the Risk signals panel.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: richer feedback, but a heavier recurring task.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Lightweight popover: Accurate/Inaccurate feedback reduced response effort.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: faster interaction, but it retained the same placement assumption.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

VALIDATION

Testing overturned my initial placement assumption

During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.

Within that workflow, the study validated both Accurate / Inaccurate feedback and additional-signal classification.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

“I wasn’t able to find how to give feedback on the flagged signal.”

“I wasn’t able to find how to give feedback on the flagged signal.”

I initially reused a hover tooltip because it matched an existing feedback pattern.

UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.
UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.

“I always go to the justification… it helps me clarify flagged risk content.”

“I always go to the justification… it helps me clarify flagged risk content.”

“I always go to the justification… it helps me clarify flagged risk content.”

Officers consistently opened the justification to understand why content had been flagged before responding.

This revealed that feedback had to follow their judgment path.

Officers consistently opened the justification to understand why content had been flagged before responding. This revealed that feedback had to follow their judgment path.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.
UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

DESIGN DECISION

I moved feedback to the judgment moment

I embedded quick signal validation inside Risk alert justifications. Then I reorganized the component around the officer’s decision sequence.

Screenshot of the Behavox alert panel showing the Risk scenario row with a  "View justification" link highlighted. An arrow points to it from a caption  reading "Where users actually went," showing the decision point where users  naturally looked for context before taking action.
Screenshot of the Behavox alert panel showing the Risk scenario row with a  "View justification" link highlighted. An arrow points to it from a caption  reading "Where users actually went," showing the decision point where users  naturally looked for context before taking action.
Before:

Before: All fields carried equal weight, nothing signals what matters most.

Before: All fields carry equal weight, nothing signals what matters most.

All fields carry equal weight, nothing signals what matters most.

1

2

The flagged signal competed with secondary data fields

2

1

Raw regulatory metadata interrupted the decision path

Improved risk alert justification component in collapsed state.  Flagged compliance signal elevated with orange left border.  Risk scenario and investigation status surfaced inline.  Supporting evidence hidden by default to reduce cognitive load.  UX case study by Yanick, senior product designer.
Improved risk alert justification component in collapsed state.  Flagged compliance signal elevated with orange left border.  Risk scenario and investigation status surfaced inline.  Supporting evidence hidden by default to reduce cognitive load.  UX case study by Yanick, senior product designer.
After:

After: I moved the feedback loop to the moment where the most useful judgment was already happening.

I moved the feedback loop to the moment where the most useful judgment was already happening.

1

1

The flagged signal led the hierarchy

2

2

Officers could mark risk signals without leaving the investigation

Improved risk alert justification component showing expanded  Supporting evidence panel. Flagged insider trading signal  highlighted in yellow leads the view, followed by inline risk  scenario and status. Secondary regulatory metadata accessible  via clean external link. UX case study by Yanick, senior  product designer.
Improved risk alert justification component showing expanded  Supporting evidence panel. Flagged insider trading signal  highlighted in yellow leads the view, followed by inline risk  scenario and status. Secondary regulatory metadata accessible  via clean external link. UX case study by Yanick, senior  product designer.

3

3

Supporting evidence used progressive disclosure

4

4

Secondary metadata moved out of the primary review path

Secondary regulatory metadata moved out of the primary review path

Secondary regulatory metadata moved out of the primary review path

DELIVERY

The validated officer path launched first

Data aggregation limits blocked the dashboard launch. To resolve this, I negotiated a phased delivery with the PM rather than waiting.

Data aggregation limits blocked the dashboard launch. To resolve this, I negotiated a phased delivery with the PM rather than waiting.

A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.
A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.

Planned for Phase 2

Deferred

This allowed engineering to work on supporting reliable manager oversight.

Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.

Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.

COMPLEXITY

We measured speed without ignoring oversight risk

The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.

A UI card titled "Speed + Adoption" outlining key performance indicators for the feedback path, including feedback completion, review friction, and alert accuracy.
A UI card titled "Speed + Adoption" outlining key performance indicators for the feedback path, including feedback completion, review friction, and alert accuracy.
A UI card titled "Quality + Oversight" outlining key metrics for system health, specifically measuring the reduction of false positives and monitoring manager visibility.
A UI card titled "Quality + Oversight" outlining key metrics for system health, specifically measuring the reduction of false positives and monitoring manager visibility.

MEASUREMENT

We measured speed without ignoring oversight risk

The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.

MEASUREMENT

We measured speed without ignoring oversight risk

The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.

MEASUREMENT

We measured speed without ignoring oversight risk

The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.

A UI card titled "Speed + Adoption" outlining key performance indicators for the feedback path, including feedback completion, review friction, and alert accuracy.
A UI card titled "Speed + Adoption" outlining key performance indicators for the feedback path, including feedback completion, review friction, and alert accuracy.
A UI card titled "Quality + Oversight" outlining key metrics for system health, specifically measuring the reduction of false positives and monitoring manager visibility.
A UI card titled "Quality + Oversight" outlining key metrics for system health, specifically measuring the reduction of false positives and monitoring manager visibility.
A UI card titled "Quality + Oversight" outlining key metrics for system health, specifically measuring the reduction of false positives and monitoring manager visibility.

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox's UI.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox's UI.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox's UI.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

LESSONS

LESSONS

Feedback systems work only when they meet real judgment behavior

1

Behavior beats familiar patterns

Behavior beats familiar patterns

Compliance oficers ignored the tooltip as they made decisions in the justification flow.

Officers ignored the tooltip because decisions happened in the justification flow.

Compliance officers ignored the tooltip as they made decisions in the justification flow.

2

Feedback belongs at the judgment moment

Feedback belongs at the judgment moment

Feedback quality improved once officers could act where they reviewed evidence.

Feedback quality improved once officers could act where they reviewed evidence.

3

Phasing protects delivery momentum

Phasing protects delivery momentum

The tightest feedback path gave us the clearest signal before scaling oversight.

The tightest feedback path gave us the clearest signal before scaling oversight.

Want the full story?

This case study is the high-level view. Happy to go deeper in conversation.

Want the full story?

This case study is the high-level view. Happy to go deeper in conversation.

Want the full story?

This case study is the high-level view. Happy to go deeper in conversation.

Want the full story?

This case study is the high-level view. Happy to go deeper in conversation.

DELIVERY

We sequenced feedback before manager oversight

Data aggregation limits blocked the dashboard launch. I negotiated a phased delivery with the PM rather than waiting.

This allowed engineering to work on supporting reliable manager oversight.

A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.

Deferred for phase 2

A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.

Deferred for phase 2

Up next

PROBLEM

Officer judgment lived outside the tuning path

Feedback required a separate step after review, delaying when analysts could use the signal.

The handoff separated evidence, judgment, and tuning. Managers also lacked a reliable aggregate view of review coverage and model health.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.
UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

“I wasn’t able to find how to give feedback on the flagged signal.”

I initially reused a hover tooltip because it matched an existing feedback pattern.

UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.
UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.

“I always go to the justification… it helps me clarify flagged risk content.”

Officers consistently opened the justification to understand why content had been flagged before responding.

VALIDATION

Testing overturned my initial placement assumption

During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.

Within that workflow, the study validated both Accurate / Inaccurate feedback and additional-signal classification.

LESSONS

What this taught me about feedback loops, behavior, and scale

1

Behavior beats familiar patterns

Compliance officers ignored the tooltip as they made decisions in the justification flow.

2

Feedback belongs at judgment

Feedback quality improved once officers could act where they reviewed evidence.

3

Phasing protects momentum

Shipping the tightest feedback path gave us the clearest signal before scaling oversight.

Turning risk alert reviews into an ML feedback loop

How I moved frontline compliance judgment into the alert-review moment while Behavox analysts retained control of model tuning.

Compliance officers had the clearest context to judge whether an alert was accurate, but their judgment sat outside the tuning path.

I designed an in-context feedback path that captured their assessment beside the alert evidence and sent it to Behavox analysts.

IMPACT

64%

Boost in risk alert accuracy

31%

Increase in usability score

Improved ML training loop

My role and scope

Product Designer: Working with product a manager, a researcher, and two engineers

Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery

Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM

A code snippet illustration showing a configuration screen with Python and pseudo-SQL logic, highlighting "tuning rules" for a machine learning model, with an icon of a person looking confused by the logic.

Opaque tuning process

Users described the flow as slow and confusing. They could not see how feedback improved the model.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.

Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.

Legacy tuning path:

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.

Legacy tuning path:

Product design visualization showing three compliance roles: Compliance Officer, Compliance Manager, and Behavox Analyst. The Compliance Manager card is highlighted to show an overlooked role in the ML model monitoring/ process.”

Oversight needs were missing from the feedback loop

Compliance risk managers lacked oversight tools. This hurt model health monitoring.

COMPLEXITY

The solution had to shorten tuning without weakening oversight

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."
Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

Thanks for reading.

Thanks for reading.

Thanks for reading.

Thanks for reading.

You're

BEHIND THE SOLUTION

The visible change was small. The sequencing decision was the real work.

The harder question was what to ship first across three roles with different responsibilities and data needs.

Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.

Create a free website with Framer, the website builder loved by startups, designers and agencies.