BEHAVOX | AI-POWERED COMPLIANCE

Turning risk alert reviews into an ML feedback loop

Compliance officers had the context to judge whether an alert was accurate, but their judgment sat outside the ML model-tuning path. I designed an in-context feedback flow capturing their input beside the alert evidence. Meanwhile, Behavox analysts kept control of ML changes.

Role: Lead product designer, partnering with a PM, a researcher, and two engineers

Scope: Alert-review feedback workflow, signal validation and classification, and deferred manager oversight

Contribution: Led workflow and interaction design; secured a two-week validation window and recommended phased delivery.

64%

Increase in risk alert accuracy

31%

Improvement in usability score

SOLUTION

Compliance officers gained a direct feedback path

Officers gained a direct feedback path

The launched workflow kept feedback inside alert review. It connected frontline judgment to analyst tuning input without transferring control over model changes.

Validate detected signals

Validate detected signals

A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.

1

1

1

Evidence stays in view

Evidence stays in view

2

2

2

Feedback sits beside the decision context

Feedback sits beside the decision context

Classify additional signals

Classify additional signals

End-to-end workflow showing highlighted text linked to a compliance scenario via an inline menu and a side panel form.
Inline context menu on highlighted email text with options to create a label or link the content to a compliance scenario.
Side panel form for linking highlighted content to a compliance scenario, with dropdowns for scenario and use case and save actions.
Side panel form for linking highlighted content to a compliance scenario, with dropdowns for scenario and use case and save actions.

1

1

Classification starts from selected evidence

Classification starts from selected evidence

2

2

New signals link to existing scenarios

New signals link to existing scenarios

1

Classification starts from selected evidence

2

New signals link to existing scenarios

Manager oversight stayed separate from the launch

The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.

The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.

A full compliance dashboard view in dark mode, showing a sidebar menu on the left, a list of various flagged risk alerts in the center, and an open "Risk alert justifications" panel on the right with a chat transcript between Mike Harris and Cheryl Smith.
A UI component labeled "Risk alert justifications" displaying a flagged message: "I can send you a $500 thank-you if you approve the renewal before Friday," identified as an indicator of an "Improper Inducement – Offer to Bribe" risk, with buttons to validate the signal as accurate or inaccurate.
An expanded UI component showing "Risk alert justifications," detailing the signal reason, risk type, policy reference, and specific metadata like matched phrase and participant count for a flagged "Improper Inducement – Offer to Bribe" risk.

The documented concept covered alert accuracy, validation coverage, review activity, and trend-level monitoring for later delivery.

SOLUTION

The ML feedback loop now starts inside the review flow

I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.

SOLUTION

The ML feedback loop now starts inside the review flow

I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.

BEHIND THE SOLUTION

The feedback loop had to fit how officers worked

BEHIND THE SOLUTION

The feedback loop had to fit how officers worked

BEHIND THE SOLUTION

The feedback loop had to fit how officers worked

The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.

The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.

PROBLEM

Officer judgment lived outside the tuning path

PROBLEM

Officer judgment lived outside the tuning path

PROBLEM

Officer judgment lived outside the tuning path

The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.

The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.
UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise. Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.

Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

CHALLENGE

Three roles created one sequencing constraint

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."
Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

EXPLORATION

I evaluated new ways to capture officer judgment

Three key concepts varied placement and response depth before testing moved feedback into risk alert justifications.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Structured popover: guided questions attached directly to selected evidence.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Structured popover: guided questions attached directly to selected evidence.

Tradeoff: strong context, but feedback appeared before officers' moment.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: strong context, but feedback appeared before officers' moment.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Guided signal review: structured input inside the Risk signals panel.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Guided signal review: structured input inside the Risk signals panel.

Tradeoff: richer feedback, but a heavier recurring task.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: richer feedback, but a heavier recurring task.

In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.
In-context risk alert review modal allowing compliance officers to confirm or reject an ML bribery signal directly within highlighted text.

Lightweight popover: Accurate/Inaccurate feedback reduced response effort.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Lightweight popover: Accurate/Inaccurate feedback reduced response effort.

Tradeoff: faster interaction, but it retained the same placement assumption.

Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.

Tradeoff: faster interaction, but it retained the same placement assumption.

VALIDATION

Testing overturned my initial placement assumption

Our 2-week unmoderated study using Maze revealed that officers looked to the justification before giving feedback.

With the project at risk of being paused, I partnered with the product manager to make the case for a two-week testing window. We secured time to test the officer workflow before proceeding.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

“I wasn’t able to find how to give feedback on the flagged signal.”

I initially reused a hover tooltip because it matched an existing feedback pattern.

UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.
UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.

“I always go to the justification… it helps me clarify flagged risk content.”

“I always go to the justification… it helps me clarify flagged risk content.”

Officers consistently opened the justification to understand why content had been flagged before responding.

This revealed that feedback had to follow their judgment path.

Officers consistently opened the justification to understand why content had been flagged before responding. This revealed that feedback had to follow their judgment path.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.
UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

DESIGN DECISION

I moved feedback to the judgment moment

I embedded quick signal validation inside Risk alert justifications. Then I reorganized the component around the officer’s decision sequence.

Screenshot of the Behavox alert panel showing the Risk scenario row with a  "View justification" link highlighted. An arrow points to it from a caption  reading "Where users actually went," showing the decision point where users  naturally looked for context before taking action.
Screenshot of the Behavox alert panel showing the Risk scenario row with a  "View justification" link highlighted. An arrow points to it from a caption  reading "Where users actually went," showing the decision point where users  naturally looked for context before taking action.
Before:

Before: All fields carried equal weight, nothing signals what matters most.

Before: All fields carry equal weight, nothing signals what matters most.

All fields carry equal weight, nothing signals what matters most.

1

1

2

The flagged signal competed with secondary data fields

The flagged signal competed with secondary data fields

2

2

1

Raw regulatory metadata interrupted the decision path

Raw regulatory metadata interrupted the decision path

Improved risk alert justification component in collapsed state.  Flagged compliance signal elevated with orange left border.  Risk scenario and investigation status surfaced inline.  Supporting evidence hidden by default to reduce cognitive load.  UX case study by Yanick, senior product designer.
Improved risk alert justification component in collapsed state.  Flagged compliance signal elevated with orange left border.  Risk scenario and investigation status surfaced inline.  Supporting evidence hidden by default to reduce cognitive load.  UX case study by Yanick, senior product designer.
After:

After: I moved the feedback loop to the moment where the most useful judgment was already happening.

I moved the feedback loop to the moment where the most useful judgment was already happening.

1

1

1

The flagged signal led the hierarchy

The flagged signal led the hierarchy

2

2

2

Officers could mark risk signals without leaving the investigation

Officers could mark risk signals without leaving the investigation

Improved risk alert justification component showing expanded  Supporting evidence panel. Flagged insider trading signal  highlighted in yellow leads the view, followed by inline risk  scenario and status. Secondary regulatory metadata accessible  via clean external link. UX case study by Yanick, senior  product designer.
Improved risk alert justification component showing expanded  Supporting evidence panel. Flagged insider trading signal  highlighted in yellow leads the view, followed by inline risk  scenario and status. Secondary regulatory metadata accessible  via clean external link. UX case study by Yanick, senior  product designer.

3

3

3

Supporting evidence used progressive disclosure

Supporting evidence used progressive disclosure

4

4

4

Secondary metadata moved out of the primary review path

Secondary regulatory metadata moved out of the primary review path

Secondary regulatory metadata moved out of the primary review path

Secondary regulatory metadata moved out of the primary review path

DELIVERY

The validated officer path launched first

Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.

Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.

A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.
A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.

Planned for Phase 2

Deferred

Analysts received officer feedback sooner, while managers waited for aggregate oversight.

Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.

COMPLEXITY

We measured speed without ignoring oversight risk

The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.

Phase 1 immediate proof metrics: feedback completion, review friction, participation across reviewed alerts, risk alert accuracy, and false-positive movement.
Phase 1 immediate proof metrics: feedback completion, review friction, participation across reviewed alerts, risk alert accuracy, and false-positive movement.
Phase 2 oversight guardrails: manager visibility, validation coverage, and trend-level monitoring, activated once data aggregation supported reliable oversight.
Phase 2 oversight guardrails: manager visibility, validation coverage, and trend-level monitoring, activated once data aggregation supported reliable oversight.

MEASUREMENT

Measurement followed the phased product state

The measurement framework covered Phase 1 adoption, friction, and signal quality. Meanwhile manager oversight measures proposed for Phase 2.

MEASUREMENT

Measurement followed the phased product state

The measurement framework covered Phase 1 adoption, friction, and signal quality. Meanwhile manager oversight measures proposed for Phase 2.

Phase 1 immediate proof metrics: feedback completion, review friction, participation across reviewed alerts, risk alert accuracy, and false-positive movement.
Phase 1 immediate proof metrics: feedback completion, review friction, participation across reviewed alerts, risk alert accuracy, and false-positive movement.
Phase 2 oversight guardrails: manager visibility, validation coverage, and trend-level monitoring, activated once data aggregation supported reliable oversight.
Phase 2 oversight guardrails: manager visibility, validation coverage, and trend-level monitoring, activated once data aggregation supported reliable oversight.

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

“Yanick pushed our products forward in terms of design. His general ingenuity had a significant impact on Behavox.“

P{rofile photo of Gustavo Pelaez, Senior Product Designer

Artsiom Mezin

Sr. Engineering Manager

KEY LEARNING

KEY LEARNING

Feedback capture and monitoring need different release criteria

Officer feedback launched while the manager dashboard waited for reliable aggregate data. I now assess dependencies for each capability before deciding what must ship together.

Want the full story?

Happy to go deeper in conversation.

Want the full story?

Happy to go deeper in conversation.

Want the full story?

Happy to go deeper in conversation.

Want the full story?

Happy to go deeper in conversation.

Up next

PROBLEM

Officer judgment lived outside the tuning path

The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.

The handoff separated evidence, judgment, and tuning. Managers also lacked a reliable aggregate view of review coverage and model health.

UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.
UI wireframe demonstrating "Where I expected feedback," showing a pop-up modal overlay titled "Review risk alert signal" with "Accurate" and "Inaccurate" buttons anchored above a highlighted chat message from Oliver Torres.

I initially reused a hover tooltip because it matched an existing feedback pattern.

“I wasn’t able to find how to give feedback on the flagged signal.”

UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.
UI wireframe illustrating "Where users went," showing a compliance interface with a yellow risk scenario banner for "IDB-Insider trading" and a "View justification" dropdown link above the chat details.

Officers consistently opened the justification to understand why content had been flagged before responding.

“I always go to the justification… it helps me clarify flagged risk content.”

VALIDATION

Testing overturned my initial placement assumption

During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.

With the project at risk of being paused, I partnered with the product manager to make the case for a two-week testing window. We secured time to test the officer workflow before proceeding.

Turning risk alert reviews into an ML feedback loop

BEHAVOX | AI-POWERED COMPLIANCE

Compliance officers had the context to judge whether an alert was accurate, but their judgment sat outside the ML model-tuning path. I designed an in-context feedback flow capturing their input beside the alert evidence. Meanwhile, Behavox analysts kept control of ML changes.

I designed an in-context feedback flow that captured their assessment beside the alert evidence. Meanwhile, Behavox analysts retained control of model changes.

IMPACT

64%

Increase in risk alert accuracy

31%

Improvement in usability score

In-context feedback path

My role and scope

Role: Lead product designer, partnering with a PM, a researcher, and two engineers

Scope: Compliance officer feedback workflow, manager oversight concept, evaluation, and phased delivery plan

Contribution: Led workflow and interaction design; secured a two-week validation window and recommended phased delivery

A code snippet illustration showing a configuration screen with Python and pseudo-SQL logic, highlighting "tuning rules" for a machine learning model, with an icon of a person looking confused by the logic.

Opaque tuning process

Users described the flow as slow and confusing. They could not see how feedback improved the model.

Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.

Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.
UX flow diagram showing the compliance workflow for ML risk review at Behavox. Highlights a bottleneck between Compliance Officer and Behavox Analyst during model fine-tuning, illustrating inefficiencies in the feedback loop.
Product design visualization showing three compliance roles: Compliance Officer, Compliance Manager, and Behavox Analyst. The Compliance Manager card is highlighted to show an overlooked role in the ML model monitoring/ process.”

Oversight needs were missing from the feedback loop

Compliance risk managers lacked oversight tools. This hurt model health monitoring.

COMPLEXITY

The solution had to shorten tuning without weakening oversight

Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.

Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."
Image Description: A vertical stack of three cards on a dark background, illustrating the roles mentioned in the challenge text:  Top Card: Features a portrait of a female Compliance Officer on the left, next to the text: "COMPLIANCE OFFICER. Judgment in context. Needs alert evidence at the moment of review."  Middle Card: Features a portrait of a male Behavox Analyst on the left, next to the text: "BEHAVOX ANALYST. ML tuning control. Requires control over ML model changes."  Bottom Card: Features a portrait of a female Compliance Manager on the left, next to the text: "COMPLIANCE MANAGER. Aggregate oversight. Depends on reliable aggregate data insights."

Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.

Thanks for reading.

Thanks for reading.

Thanks for reading.

Thanks for reading.

BEHIND THE SOLUTION

The feedback loop had to fit how officers worked

The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.

Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.

DELIVERY

The validated officer path launched first

Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.

Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.

A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.
A view of the manager monitoring dashboard, showcasing data on alert accuracy, weekly signal detection, scenario validation coverage, and review activity, with a label indicating that specific oversight features were deferred to Phase 2.

Planned for Phase 2

Analysts received officer feedback sooner, while managers waited for aggregate oversight.

Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.

Create a free website with Framer, the website builder loved by startups, designers and agencies.