Turning risk alert reviews into an ML feedback loop






Compliance officers could judge whether a risk alert was accurate, useful, or important. But the product gave them no way to send that judgment to behavioral analysts. Without that signal, Behavox had a weaker path to improve model quality and reduce costly alert noise.
I designed and launched a ML risk alert feedback path inside the alert review flow. It captured frontline judgment for analysts while preserving their control over model tuning.
My role and scope
My role and scope
Senior Product Designer • Team: product manager, researcher, and two engineers
Senior Product Designer • Team: product manager, researcher, and two engineers
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM
64%
Boost in risk alert accuracy
Boost in risk alert accuracy
31%
Increase in usability score
Increase in usability score
Reduced analyst dependency
Reduced analyst dependency
SOLUTION
Compliance officers gained a direct feedback path
The launched workflow let officers send their assessment to analysts from inside the alert review.


Launched officer workflow
1
1
Evidence stays in view
Evidence stays in view
2
2
Feedback sits beside the decision context
Feedback sits beside the decision context
Testing validated the embedded direction
Six compliance officers validated the embedded workflow in an unmoderated Maze study.
Six compliance officers validated the embedded workflow in an unmoderated Maze study.


Supporting evidence remains available without dominating review.
Supporting evidence remains available without dominating review.


Validation now happens where judgment forms.
Validation now happens where judgment forms.
Manager oversight remained a deliberate second phase
Manager oversight remained a deliberate second phase
Phase 2 concept (not launched). The manager dashboard concept established a future oversight direction. Yet, reliable aggregation depended on data-pipeline work outside the initial release.
Phase 2 concept (not launched). The manager dashboard concept established a future oversight direction. Yet, reliable aggregation depended on data-pipeline work outside the initial release.



SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
BEHIND THE SOLUTION
The simple review action depended on a harder sequencing decision
BEHIND THE SOLUTION
The simple review action depended on a harder sequencing decision
BEHIND THE SOLUTION
The simple review action depended on a harder sequencing decision
Officers had the clearest contextual signal, but Behavox analysts still controlled model tuning. Meanwhile, managers needed reliable oversight. The product decision was how to shorten that loop. We needed to achieve this without presenting incomplete monitoring data as trustworthy.
That created queues and slowed signal correction. Weaker alerts stayed in circulation longer than necessary.
That created queues and slowed signal correction. Weaker alerts stayed in circulation longer than necessary.
Compliance officers saw the evidence first, but analysts owned most tuning work
PROBLEM
The strongest feedback signal lived outside the tuning path
PROBLEM
The strongest feedback signal lived outside the tuning path
PROBLEM
The strongest feedback signal lived outside the tuning path
Compliance officers developed the clearest contextual judgment while investigating alerts.
Yet, submitting that judgment required a separate step, delaying when analysts could use it to evaluate and tune model behavior.
Yet, submitting that judgment required a separate step, delaying when analysts could use it to evaluate and tune model behavior.
Compliance officers developed the clearest contextual judgment while investigating alerts.


Stalled training loops
Review and tuning sat in queues. Real threats could be missed.
Review and tuning sat in queues. Real threats could be missed.
High operational costs
No clear visibility
Delayed risk detection
Oversight needs were missing from the feedback loop
Compliance risk managers lacked oversight tools. This hurt model health monitoring.
Compliance risk managers lacked oversight tools. This hurt model health monitoring.



CHALLENGE
The solution had to shorten tuning without weakening oversight
The new flow had to let officers act in context without breaking analyst oversight, model quality, or delivery speed.
CHALLENGE
The solution had to shorten tuning without weakening oversight
The new flow had to let officers act in context without breaking analyst oversight, model quality, or delivery speed.
CHALLENGE
The solution had to shorten tuning without weakening oversight
The new flow had to let officers act in context without breaking analyst oversight, model quality, or delivery speed.
Reduce analyst dependency
Officers had signal context. Analysts still owned tuning.
Officers had signal context. Analysts still owned tuning.
Preserve oversight
Managers needed visibility into review and tuning quality.
Managers needed visibility into review and tuning quality.
Fit data constraints
Pipeline limits shaped what could ship first.
Pipeline limitations shaped what we could ship first.
OPTIONS
The strongest path moved feedback into the review decision
I compared each concept against four criteria. I looked at feedback speed, build effort, analyst dependency, and data oversight.
The strongest option moved feedback into the review moment while keeping oversight available for a later phase.



Train the model from the alert. Officers could confirm or reject signals while reviewing evidence.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Train the model from the alert. Officers could confirm or reject signals while reviewing evidence.



Classify signals in context. Reviewers could tag new signals without leaving the investigation path.
Classify signals in context. Reviewers could tag new signals without leaving the investigation path.







Track model health at scale. Managers could monitor coverage, contributions, and tuning quality.
Track model health at scale. Managers could monitor coverage, contributions, and tuning quality.
TESTING
I made the wrong call: I optimized for a familiar pattern over actual review behavior
I reused a hover tooltip because it matched an existing feedback pattern.
Testing showed officers made decisions in the justification area, not on highlighted text.

“I wasn’t able to find how to give feedback on flagged risk signals.”
“I wasn’t able to find how to give feedback on flagged risk signals.”

I matched a familiar pattern, but missed the decision moment.
I matched a familiar pattern, but missed the decision moment.



“I always go to the justification… it helps me clarify flagged risk content.”
“I always go to the justification… it helps me clarify flagged risk content.”


The issue was not discoverability alone. Feedback was outside the place where officers formed judgment.


SOLUTION
I moved ML feedback into the risk signal justification moment
Testing revealed the gap. Officers went to the justification first, every time. So I moved feedback entry there.


Before:
Before:
Before:
1
1
Flagged signal lacks emphasis, buried among secondary fields.
2
2
Regulatory data shows raw URL, interrupts the decision flow.
3
All fields carry equal weight, nothing signals what matters most.
4
Competing and hidden secondary metadata.
3
All fields carry equal weight, nothing signals what matters most.
4
Competing and hidden secondary metadata.




After:
After:
After:
1
1
Lead with the flagged signal.
2
3
Nested feedback mechanism
3
2
Surface only the information needed to judge the signal.


4
4
Hide secondary metadata until needed.
5
5
Keep regulatory evidence accessible without disrupting review.
Officers could address the signal while keeping secondary metadata available but out of the way.
PIVOT
We sequenced feedback before manager oversight

Planned for Phase 2
Data aggregation limits made the dashboard too risky for launch.

Deferred

Deferred
To maintain momentum, I secured a two-week validation window for the officer feedback path.
To maintain momentum, I secured a two-week validation window for the officer feedback path.
COMPLEXITY
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.




HANDOFF
I converted the validated flow into a phased rollout
The officer feedback path was prepared for implementation first.
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
I documented the final states, measurement plan, and manager oversight path for Phase 2.
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.



Sequencing the officer path first protected delivery momentum without forcing an unreliable oversight experience.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
HANDOFF
I turned the validated flow into a delivery-ready handoff
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
HANDOFF
I turned the validated flow into a delivery-ready handoff
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.





LESSONS
LESSONS
Feedback systems work only when they meet real judgment behavior
1
Behavior beats familiar patterns
Behavior beats familiar patterns
Compliance oficers ignored the tooltip as they made decisions in the justification flow.
Officers ignored the tooltip because decisions happened in the justification flow.
Compliance officers ignored the tooltip as they made decisions in the justification flow.
2
Feedback belongs at the judgment moment
Feedback belongs at the judgment moment
Feedback quality improved once officers could act where they reviewed evidence.
Feedback quality improved once officers could act where they reviewed evidence.
3
Phasing protects delivery momentum
Phasing protects delivery momentum
The tightest feedback path gave us the clearest signal before scaling oversight.
The tightest feedback path gave us the clearest signal before scaling oversight.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
PIVOT
We sequenced feedback before manager oversight
Data aggregation limits made the dashboard too risky for launch.
To maintain momentum, I secured a two-week validation window for the officer feedback path.

Deferred for phase 2

Deferred for phase 2
BEHIND THE SOLUTION
The strongest feedback signal lived outside the tuning path
Compliance officers developed the clearest contextual judgment while investigating alerts.
Yet, submitting that judgment required a separate step, delaying when analysts could use it to evaluate and tune model behavior.


“I wasn’t able to find how to give feedback on flagged risk signals.”



“I always go to the justification… it helps me clarify flagged risk content.”

I matched a familiar pattern, but missed the decision moment. The issue was not discoverability alone. Feedback was outside the place where officers formed judgment.
Testing
I made the wrong call: I optimized for a familiar pattern over actual review behavior
I reused a hover tooltip for ML feedback because it matched an existing pattern. Testing showed officers made decisions in the justification area, not on highlighted text.
LESSONS
What this taught me about feedback loops, behavior, and scale
1
Behavior beats familiar patterns
Compliance officers ignored the tooltip as they made decisions in the justification flow.
2
Feedback belongs at judgment
Feedback quality improved once officers could act where they reviewed evidence.
3
Phasing protects momentum
Shipping the tightest feedback path gave us the clearest signal before scaling oversight.
Turning risk alert reviews into an ML feedback loop

Compliance officers could judge whether a risk alert was accurate, useful, or important. But the product gave them no way to send that judgment to behavioral analysts.
Without that signal, Behavox had a weaker path to improve model quality and reduce costly alert noise.
I designed and launched a ML risk alert feedback path inside the alert review flow. It captured frontline judgment for analysts while preserving their control over model tuning.
IMPACT
64%
Boost in risk alert accuracy
31%
Increase in usability score
Reduced analyst dependency
My role and scope
Senior Product Designer • Team: product manager, researcher, and two engineers
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM

Opaque tuning process
Users described the flow as slow and confusing. They could not see how feedback improved the model.
Stalled training loops
Review and tuning sat in queues. Real threats could be missed.
High operational costs
No clear visibility
Delayed risk detection




Oversight needs were missing from the feedback loop
Compliance risk managers lacked oversight tools. This hurt model health monitoring.
COMPLEXITY
The solution had to shorten tuning without weakening oversight
Our new model had to work across onboarding, referral, travel rewards, and future incentives.
Reduce dependency
Officers had signal context. Analysts still owned tuning.
Preserve oversight
Managers needed visibility into review and tuning quality.
Fit data constraints
Pipeline limitations shaped what we could ship first.
You're


