BEHAVOX | AI-POWERED COMPLIANCE
Turning risk alert reviews into an ML feedback loop






Compliance officers had the context to judge whether an alert was accurate, but their judgment sat outside the ML model-tuning path. I designed an in-context feedback flow capturing their input beside the alert evidence. Meanwhile, Behavox analysts kept control of ML changes.
Role: Lead product designer, partnering with a PM, a researcher, and two engineers
Scope: Alert-review feedback workflow, signal validation and classification, and deferred manager oversight
Contribution: Led workflow and interaction design; secured a two-week validation window and recommended phased delivery.
64%
Increase in risk alert accuracy
31%
Improvement in usability score
SOLUTION
Compliance officers gained a direct feedback path
Officers gained a direct feedback path
The launched workflow kept feedback inside alert review. It connected frontline judgment to analyst tuning input without transferring control over model changes.
Validate detected signals
Validate detected signals




1
1
1
Evidence stays in view
Evidence stays in view
2
2
2
Feedback sits beside the decision context
Feedback sits beside the decision context
Classify additional signals
Classify additional signals




1
1
Classification starts from selected evidence
Classification starts from selected evidence
2
2
New signals link to existing scenarios
New signals link to existing scenarios
1
Classification starts from selected evidence
2
New signals link to existing scenarios
Manager oversight stayed separate from the launch
The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.
The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.



The documented concept covered alert accuracy, validation coverage, review activity, and trend-level monitoring for later delivery.
SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
BEHIND THE SOLUTION
The feedback loop had to fit how officers worked
BEHIND THE SOLUTION
The feedback loop had to fit how officers worked
BEHIND THE SOLUTION
The feedback loop had to fit how officers worked
The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.
The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.
PROBLEM
Officer judgment lived outside the tuning path
PROBLEM
Officer judgment lived outside the tuning path
PROBLEM
Officer judgment lived outside the tuning path
The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.
The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.


Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise. Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.
Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.


Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
EXPLORATION
I evaluated new ways to capture officer judgment
Three key concepts varied placement and response depth before testing moved feedback into risk alert justifications.


Structured popover: guided questions attached directly to selected evidence.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Structured popover: guided questions attached directly to selected evidence.
Tradeoff: strong context, but feedback appeared before officers' moment.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: strong context, but feedback appeared before officers' moment.


Guided signal review: structured input inside the Risk signals panel.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Guided signal review: structured input inside the Risk signals panel.
Tradeoff: richer feedback, but a heavier recurring task.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: richer feedback, but a heavier recurring task.


Lightweight popover: Accurate/Inaccurate feedback reduced response effort.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Lightweight popover: Accurate/Inaccurate feedback reduced response effort.
Tradeoff: faster interaction, but it retained the same placement assumption.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: faster interaction, but it retained the same placement assumption.
VALIDATION
Testing overturned my initial placement assumption
Our 2-week unmoderated study using Maze revealed that officers looked to the justification before giving feedback.
With the project at risk of being paused, I partnered with the product manager to make the case for a two-week testing window. We secured time to test the officer workflow before proceeding.

“I wasn’t able to find how to give feedback on the flagged signal.”
I initially reused a hover tooltip because it matched an existing feedback pattern.


“I always go to the justification… it helps me clarify flagged risk content.”
“I always go to the justification… it helps me clarify flagged risk content.”
Officers consistently opened the justification to understand why content had been flagged before responding.
This revealed that feedback had to follow their judgment path.
Officers consistently opened the justification to understand why content had been flagged before responding. This revealed that feedback had to follow their judgment path.


DESIGN DECISION
I moved feedback to the judgment moment
I embedded quick signal validation inside Risk alert justifications. Then I reorganized the component around the officer’s decision sequence.


Before:
Before: All fields carried equal weight, nothing signals what matters most.
Before: All fields carry equal weight, nothing signals what matters most.
All fields carry equal weight, nothing signals what matters most.
1
1
2
The flagged signal competed with secondary data fields
The flagged signal competed with secondary data fields
2
2
1
Raw regulatory metadata interrupted the decision path
Raw regulatory metadata interrupted the decision path


After:
After: I moved the feedback loop to the moment where the most useful judgment was already happening.
I moved the feedback loop to the moment where the most useful judgment was already happening.
1
1
1
The flagged signal led the hierarchy
The flagged signal led the hierarchy
2
2
2
Officers could mark risk signals without leaving the investigation
Officers could mark risk signals without leaving the investigation


3
3
3
Supporting evidence used progressive disclosure
Supporting evidence used progressive disclosure
4
4
4
Secondary metadata moved out of the primary review path
Secondary regulatory metadata moved out of the primary review path
Secondary regulatory metadata moved out of the primary review path
Secondary regulatory metadata moved out of the primary review path
DELIVERY
The validated officer path launched first
Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.
Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.


Planned for Phase 2
Deferred
Analysts received officer feedback sooner, while managers waited for aggregate oversight.
Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.
COMPLEXITY
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.




MEASUREMENT
Measurement followed the phased product state
The measurement framework covered Phase 1 adoption, friction, and signal quality. Meanwhile manager oversight measures proposed for Phase 2.
MEASUREMENT
Measurement followed the phased product state
The measurement framework covered Phase 1 adoption, friction, and signal quality. Meanwhile manager oversight measures proposed for Phase 2.




KEY LEARNING
KEY LEARNING
Feedback capture and monitoring need different release criteria
Officer feedback launched while the manager dashboard waited for reliable aggregate data. I now assess dependencies for each capability before deciding what must ship together.
Want the full story?
Happy to go deeper in conversation.
Want the full story?
Happy to go deeper in conversation.
Want the full story?
Happy to go deeper in conversation.
Want the full story?
Happy to go deeper in conversation.
PROBLEM
Officer judgment lived outside the tuning path
The legacy tuning path required feedback as a separate step after review. This delayed when analysts could use the signal.
The handoff separated evidence, judgment, and tuning. Managers also lacked a reliable aggregate view of review coverage and model health.


I initially reused a hover tooltip because it matched an existing feedback pattern.
“I wasn’t able to find how to give feedback on the flagged signal.”


Officers consistently opened the justification to understand why content had been flagged before responding.
“I always go to the justification… it helps me clarify flagged risk content.”
VALIDATION
Testing overturned my initial placement assumption
During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.
With the project at risk of being paused, I partnered with the product manager to make the case for a two-week testing window. We secured time to test the officer workflow before proceeding.
Turning risk alert reviews into an ML feedback loop
BEHAVOX | AI-POWERED COMPLIANCE

Compliance officers had the context to judge whether an alert was accurate, but their judgment sat outside the ML model-tuning path. I designed an in-context feedback flow capturing their input beside the alert evidence. Meanwhile, Behavox analysts kept control of ML changes.
I designed an in-context feedback flow that captured their assessment beside the alert evidence. Meanwhile, Behavox analysts retained control of model changes.
IMPACT
64%
Increase in risk alert accuracy
31%
Improvement in usability score
In-context feedback path
My role and scope
Role: Lead product designer, partnering with a PM, a researcher, and two engineers
Scope: Compliance officer feedback workflow, manager oversight concept, evaluation, and phased delivery plan
Contribution: Led workflow and interaction design; secured a two-week validation window and recommended phased delivery

Opaque tuning process
Users described the flow as slow and confusing. They could not see how feedback improved the model.
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.
Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.



Oversight needs were missing from the feedback loop
Compliance risk managers lacked oversight tools. This hurt model health monitoring.
COMPLEXITY
The solution had to shorten tuning without weakening oversight
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.


Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
BEHIND THE SOLUTION
The feedback loop had to fit how officers worked
The challenge was finding where officers formed judgment. Concurrently, we needed to decide what could launch while manager oversight waited for reliable data.
Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.
DELIVERY
The validated officer path launched first
Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.
Data aggregation limits blocked the manager dashboard. With the product manager, I recommended launching the validated officer workflow first.


Planned for Phase 2
Analysts received officer feedback sooner, while managers waited for aggregate oversight.
Note: Early concepts utilized the legacy production baseline. I fully migrated the final designs to the latest design system.


