Turning risk alert reviews into an ML feedback loop






How I moved frontline compliance judgment into the alert-review moment while Behavox analysts retained control of model tuning.
Compliance officers had the clearest context to judge whether an alert was accurate, but their judgment sat outside the tuning path. I designed an in-context feedback path that captured their assessment beside the alert evidence and sent it to Behavox analysts.
My role and scope
My role and scope
Product Designer: Working with product a manager, a researcher, and two engineers
Product Designer: Working with product a manager, a researcher, and two engineers
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM
64%
Boost in risk alert accuracy
Boost in risk alert accuracy
31%
Increase in usability score
Increase in usability score
Improved ML training loop
Improved ML training loop
SOLUTION
Compliance officers gained a direct feedback path
Officers gained a direct feedback path
The launched workflow kept feedback inside alert review. It connected frontline judgment to analyst tuning input without transferring control over model changes.
Validate detected signals
Validate detected signals


1
1
Evidence stays in view
Evidence stays in view
Evidence stays in view
2
2
Feedback sits beside the decision context
Feedback sits beside the decision context
Feedback sits beside the decision context
Classify additional signals
Classify additional signals




1
Classification starts from selected evidence
Classification starts from selected evidence
2
2
New signals link to existing scenarios
New signals link to existing scenarios
1
Classification starts from selected evidence
2
New signals link to existing scenarios
40
Manager oversight stayed separate from the launch
The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.
The dashboard defined a future monitoring direction and remained explicitly separate from the launched officer workflow.



The documented concept covered alert accuracy, validation coverage, review activity, and trend-level monitoring for later delivery.
SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
SOLUTION
The ML feedback loop now starts inside the review flow
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
BEHIND THE SOLUTION
The simple review action depended on a harder sequencing decision
BEHIND THE SOLUTION
The simple review action depended on a harder sequencing decision
BEHIND THE SOLUTION
The visible change was small. The sequencing decision was the real work.
The harder question was what to ship first across three roles with different responsibilities and data needs.
Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.
Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.
The harder question was what to ship first across three roles with different responsibilities and data needs.
PROBLEM
Officer judgment lived outside the tuning path
PROBLEM
Officer judgment lived outside the tuning path
PROBLEM
Officer judgment lived outside the tuning path
Feedback required a separate step after review, delaying when analysts could use the signal.
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise, but the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it. Managers also lacked a reliable aggregate view of review coverage and model health.
Feedback required a separate step after review, delaying when analysts could use the signal.

Legacy tuning path:
Legacy tuning path:

Legacy tuning path:
Legacy tuning path:
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise. Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise.
Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
CHALLENGE
Three roles created one sequencing constraint
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.


Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
EXPLORATION
I evaluated new ways to capture officer judgment
Three key concepts varied placement and response depth before testing moved feedback into risk alert justifications.


Structured popover: guided questions attached directly to selected evidence.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: strong context, but feedback appeared before officers' moment.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.


Guided signal review: structured input inside the Risk signals panel.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: richer feedback, but a heavier recurring task.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.


Lightweight popover: Accurate/Inaccurate feedback reduced response effort.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
Tradeoff: faster interaction, but it retained the same placement assumption.
Train the model from alerts. Officers could confirm/reject signals while reviewing evidence.
VALIDATION
Testing overturned my initial placement assumption
During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.
Within that workflow, the study validated both Accurate / Inaccurate feedback and additional-signal classification.

“I wasn’t able to find how to give feedback on the flagged signal.”
“I wasn’t able to find how to give feedback on the flagged signal.”
I initially reused a hover tooltip because it matched an existing feedback pattern.


“I always go to the justification… it helps me clarify flagged risk content.”
“I always go to the justification… it helps me clarify flagged risk content.”
“I always go to the justification… it helps me clarify flagged risk content.”
Officers consistently opened the justification to understand why content had been flagged before responding.
This revealed that feedback had to follow their judgment path.
Officers consistently opened the justification to understand why content had been flagged before responding. This revealed that feedback had to follow their judgment path.


DESIGN DECISION
I moved feedback to the judgment moment
I embedded quick signal validation inside Risk alert justifications. Then I reorganized the component around the officer’s decision sequence.


Before:
Before:
Before:
All fields carry equal weight, nothing signals what matters most.
1
2
The flagged signal competed with secondary data fields
The flagged signal competed with secondary data fields
2
1
Raw regulatory metadata interrupted the decision path
Raw regulatory metadata interrupted the decision path


After:
After:
After:
1
1
The flagged signal led the hierarchy
The flagged signal led the hierarchy
2
2
Officers could mark risk signals without leaving the investigation
Officers could mark risk signals without leaving the investigation


3
3
Supporting evidence used progressive disclosure
Supporting evidence used progressive disclosure
4
4
Secondary metadata moved out of the primary review path
Secondary regulatory metadata moved out of the primary review path
Secondary regulatory metadata moved out of the primary review path
The product change was not simply a cleaner component. It moved the feedback loop to the moment where the most useful judgment was already happening.
The product change was not simply a cleaner component. It moved the feedback loop to the moment where the most useful judgment was already happening.
PIVOT
We sequenced feedback before manager oversight

Planned for Phase 2
Data aggregation limits made the dashboard too risky for launch.

Deferred

Deferred
To maintain momentum, I secured a two-week validation window for the officer feedback path.
To maintain momentum, I secured a two-week validation window for the officer feedback path.
COMPLEXITY
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.




HANDOFF
I converted the validated flow into a phased rollout
The officer feedback path was prepared for implementation first.
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
I documented the final states, measurement plan, and manager oversight path for Phase 2.
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.



Sequencing the officer path first protected delivery momentum without forcing an unreliable oversight experience.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
HANDOFF
I turned the validated flow into a delivery-ready handoff
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
HANDOFF
I turned the validated flow into a delivery-ready handoff
I handed off the officer feedback flow and documented the Phase 2 dashboard path for future rollout.
Final screens use the latest design system. Earlier concepts use the legacy system because it was the available production baseline.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.
MEASUREMENT
We measured speed without ignoring oversight risk
The goal was not only faster feedback. We also tracked whether the new path improved signal quality without weakening oversight.





LESSONS
LESSONS
Feedback systems work only when they meet real judgment behavior
1
Behavior beats familiar patterns
Behavior beats familiar patterns
Compliance oficers ignored the tooltip as they made decisions in the justification flow.
Officers ignored the tooltip because decisions happened in the justification flow.
Compliance officers ignored the tooltip as they made decisions in the justification flow.
2
Feedback belongs at the judgment moment
Feedback belongs at the judgment moment
Feedback quality improved once officers could act where they reviewed evidence.
Feedback quality improved once officers could act where they reviewed evidence.
3
Phasing protects delivery momentum
Phasing protects delivery momentum
The tightest feedback path gave us the clearest signal before scaling oversight.
The tightest feedback path gave us the clearest signal before scaling oversight.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
Want the full story?
This case study is the high-level view. Happy to go deeper in conversation.
PIVOT
We sequenced feedback before manager oversight
Data aggregation limits made the dashboard too risky for launch.
To maintain momentum, I secured a two-week validation window for the officer feedback path.

Deferred for phase 2

Deferred for phase 2
PROBLEM
Officer judgment lived outside the tuning path
Feedback required a separate step after review, delaying when analysts could use the signal.
The handoff separated evidence, judgment, and tuning. Managers also lacked a reliable aggregate view of review coverage and model health.


“I wasn’t able to find how to give feedback on the flagged signal.”
I initially reused a hover tooltip because it matched an existing feedback pattern.


“I always go to the justification… it helps me clarify flagged risk content.”
Officers consistently opened the justification to understand why content had been flagged before responding.
VALIDATION
Testing overturned my initial placement assumption
During my 2-week unmoderated Maze study, eight compliance officers validated both feedback modes inside the justification workflow.
Within that workflow, the study validated both Accurate / Inaccurate feedback and additional-signal classification.
LESSONS
What this taught me about feedback loops, behavior, and scale
1
Behavior beats familiar patterns
Compliance officers ignored the tooltip as they made decisions in the justification flow.
2
Feedback belongs at judgment
Feedback quality improved once officers could act where they reviewed evidence.
3
Phasing protects momentum
Shipping the tightest feedback path gave us the clearest signal before scaling oversight.
Turning risk alert reviews into an ML feedback loop

How I moved frontline compliance judgment into the alert-review moment while Behavox analysts retained control of model tuning.
Compliance officers had the clearest context to judge whether an alert was accurate, but their judgment sat outside the tuning path.
I designed an in-context feedback path that captured their assessment beside the alert evidence and sent it to Behavox analysts.
IMPACT
64%
Boost in risk alert accuracy
31%
Increase in usability score
Improved ML training loop
My role and scope
Product Designer: Working with product a manager, a researcher, and two engineers
Scope: Officer feedback flow, manager oversight concept, evaluation, and phased delivery
Contribution: Led the workflow and interaction design. Tested the officer direction with six compliance officers. Recommended and documented the phase split with the PM

Opaque tuning process
Users described the flow as slow and confusing. They could not see how feedback improved the model.
Alert quality shaped the review itself. Officers had the investigation context to distinguish relevant signals from noise. Yet, the separate handoff delayed when analysts could use that judgment and separated it from the evidence context that informed it.

Legacy tuning path:

Legacy tuning path:

Oversight needs were missing from the feedback loop
Compliance risk managers lacked oversight tools. This hurt model health monitoring.
COMPLEXITY
The solution had to shorten tuning without weakening oversight
Officers needed context, analysts needed control, and managers needed data the system could not yet aggregate reliably.


Any Phase 1 path had to shorten feedback without requiring unavailable aggregation or bypassing analyst control.
You're
BEHIND THE SOLUTION
The visible change was small. The sequencing decision was the real work.
The harder question was what to ship first across three roles with different responsibilities and data needs.
Officers had the context, analysts retained tuning control, and managers needed reliable aggregate oversight. The sequencing decision was how to shorten the loop without treating incomplete monitoring data as trustworthy.


