Skip to content
0xdevabirPublic

About

In an authorised-push-payment scam, the victim presses "send" themselves. Their phone, PIN and location are all genuine, so checks on the sender see a normal payment. A caller pretending to be wallet staff, a "lottery fee", an "investment" with daily returns: the money goes to a mule wallet and is cashed out at an agent within minutes.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

64 Commits

Folders and files

Repository files navigation

FraudLens

FraudLens

Stop the scam before the money moves: real-time fraud decisions for mobile money, explained in Bangla and reviewed by people

Python FastAPI Next.js Docker

LightGBM SHAP PostgreSQL Redis Streams Tailwind Tests Served = evaluated

Quick start ยท How it works ยท Results ยท Console tour ยท Three-minute demo ยท Limits


Why is this hard? In an authorised-push-payment scam, the victim presses "send" themselves. Their phone, PIN and location are all genuine, so checks on the sender see a normal payment. A caller pretending to be wallet staff, a "lottery fee", an "investment" with daily returns: the money goes to a mule wallet and is cashed out at an agent within minutes.

FraudLens answers one question, in a few milliseconds: is this payment going to a scammer, and if so, what is the gentlest thing that will stop it?

1.45M

simulated transactions
over 120 days

86.9%

of scams caught
at the warning tier

0.53%

of honest payments
interrupted

3.4 ms

p95 to score, decide
and explain, in process

0

differences between
served and evaluated
Executive summary: money kept from leaving, customers interrupted, reviewer queue and decision latency

All data is synthetic. One scam type was hidden from training on purpose, so the numbers include a scam the models had never seen. See what this is not.


๐ŸŽฏ The problem and how we answer it

The problem What FraudLens does Where to see it
The sender looks normal, because the victim is the one paying Scores the receiving wallet and its network as well as the sender Alert queue ยท Payment
A blocked payment is too late and too blunt Four tiers: allow โ†’ warn โ†’ step-up โ†’ hold, the first two decided by the customer Customer phone demo
Customers do not read English warnings Fixed, reviewed warning texts in Bangla and English Decision policy
"The model said so" is not an explanation Exact SHAP reasons as plain sentences, a rule trace, and a case note (template or optional LLM draft, checked against the evidence) Payment
Money that got through is gone A second chance at cash-out, plus victim refunds once fraud is confirmed and the mule is frozen Mule rings ยท Victim refunds
An honest customer is interrupted They can appeal; a person answers every one, and an approved hold is released (never to a frozen or confirmed-fraud wallet) Customer appeals ยท Customer phone demo
A mule hops to another provider A privacy-preserving mule consortium shares keyed tokens, not phone numbers or handsets Mule consortium
Automation should not freeze someone's money A hold waits for an analyst; a freeze needs two people Cases ยท Freeze approvals
Models go stale Verdicts become labels, challengers run in shadow mode, drift and fairness are on a dashboard Model monitoring ยท Fairness report

Hero workflow: a customer types a number and an amount. Before the money moves, the phone shows a scam warning in Bangla. A riskier payment asks them to verify again and wait out a cooling-off period. The riskiest is paused, lands at the top of an analyst's queue with its reasons, and stays paused until a person decides. If money already left, reporting the scam opens a refund claim; once two people freeze the mule and fraud is confirmed, what remains in that wallet is shared back among the victims.


๐Ÿ” How it works

One payment, start to finish

flowchart LR
    A["๐Ÿ“ฑ Customer<br/>presses send"] --> B["โšก POST /v1/score"]
    B --> C["๐Ÿงฎ Feature engine<br/>61 point-in-time features"]
    C --> D["๐Ÿง  Models<br/>transaction risk ยท mule wallet ยท anomaly"]
    D --> E{"๐Ÿ“œ Policy v2<br/>thresholds + 7 rules"}
    E -- "98.5%" --> F["โœ… Allow<br/>nothing shown"]
    E --> G["โš ๏ธ Warn<br/>scam warning in Bangla"]
    E --> H["๐Ÿ” Step-up<br/>verify again + 30 min wait"]
    E --> I["โธ๏ธ Hold<br/>paused for a person"]
    I --> J["๐Ÿง‘โ€๐Ÿ’ผ Analyst case<br/>reasons ยท network ยท verdict"]
    J --> K["๐Ÿท๏ธ Verdict becomes<br/>a training label"]
    K -.-> D
Loading

What the customer and the analyst each see

sequenceDiagram
    autonumber
    actor C as Customer
    participant W as Wallet app
    participant F as FraudLens
    actor A as Analyst
    actor S as Supervisor
    C->>W: Send เงณ9,130 to a new number
    W->>F: POST /v1/score
    F->>F: features โ†’ models โ†’ policy โ†’ reasons (a few ms)
    alt allow
        F-->>W: proceed
    else warn or step-up
        F-->>W: warning text in Bangla
        C->>W: cancel, or verify and wait
    else hold
        F-->>W: "your money is still in your wallet"
        F->>A: alert on the live feed, case opened
        A->>F: request a freeze on the receiving wallet
        S->>F: approve (must be a different person)
        A->>F: verdict: confirmed fraud or legitimate
        F-->>W: payment blocked or released
    end
Loading

The four tiers

Tier What happens Who decides False alerts per day, this tier and stricter
โœ… Allow Payment goes through, nothing is shown nobody n/a
โš ๏ธ Warn A scam warning; the customer may continue the customer 27.8
๐Ÿ” Step-up Verify again, then a 30-minute cooling-off the customer 8.4
โธ๏ธ Hold Paused, with a 30-minute review deadline a human analyst 2.3

That is out of about 5,400 scored payments a day.

The warning a customer reads before sending to a suspected mule:

เฆŸเฆพเฆ•เฆพ เฆชเฆพเฆ เฆพเฆจเง‹เฆฐ เฆ†เฆ—เง‡ เฆเฆ•เฆŸเง เฆฅเฆพเฆฎเงเฆจเฅค เฆชเงเฆฐเฆคเฆพเฆฐเฆ•เฆฐเฆพ เฆ‰เฆชเฆพเฆฏเฆผ เฆ•เฆฐเงเฆฎเฆ•เฆฐเงเฆคเฆพ, เฆฒเฆŸเฆพเฆฐเฆฟ เฆ•เฆฐเงเฆคเงƒเฆชเฆ•เงเฆท เฆฌเฆพ เฆฌเฆฟเฆจเฆฟเฆฏเฆผเง‹เฆ— เฆเฆœเง‡เฆจเงเฆŸ เฆธเง‡เฆœเง‡ เฆ เฆงเฆฐเฆจเง‡เฆฐ เฆ“เฆฏเฆผเฆพเฆฒเง‡เฆŸเง‡ เฆŸเฆพเฆ•เฆพ เฆชเฆพเฆ เฆพเฆคเง‡ เฆฌเฆฒเง‡เฅค เฆ†เฆชเฆจเฆฟ เฆ•เฆฟ เฆเฆ‡ เฆฌเงเฆฏเฆ•เงเฆคเฆฟเฆ•เง‡ เฆšเง‡เฆจเง‡เฆจ, เฆเฆฌเฆ‚ เฆจเฆฟเฆœเง‡เฆฐ เฆธเฆฟเฆฆเงเฆงเฆพเฆจเงเฆคเง‡เฆ‡ เฆŸเฆพเฆ•เฆพ เฆชเฆพเฆ เฆพเฆšเงเฆ›เง‡เฆจ?

Pause before you send. Scammers pose as upay staff, lottery officials or investment agents and ask people to send money to wallets like this one. Do you know this person, and did you decide to pay them yourself?


๐Ÿ–ฅ๏ธ Console tour

Payment page with the decision, the reasons and the case summary
Why it was flagged. Ranked reasons with the facts behind them, the rule trace, and a case summary in English or เฆฌเฆพเฆ‚เฆฒเฆพ.
A mule ring drawn as a graph of wallets linked by shared handsets and transfers
Mule rings. Wallets tied by shared handsets and transfers. Takeover victims are listed apart from members.
Customer phone demo next to the decision the platform made
Customer phone demo. Plays the wallet app against the real scoring endpoint, side by side with what FraudLens decided.
Alert queue with tier, risk score, outcome and case for each interrupted payment
Alert queue. Every interrupted payment, newest first, arriving over a live feed. Wallet numbers stay masked until a reveal is audited.
Impact simulator: a threshold slider with money saved, customers interrupted and reviewer hours
Impact simulator. Move the threshold and watch money saved, honest customers interrupted and reviewer hours move together.
Fairness report: false-alert rate by region and account age
Fairness report. Who pays for false alarms, by region, account age, balance and channel.
Page What it does
Executive summary Money stopped, customers interrupted, reviewer backlog and decision latency, from the live service
Impact simulator The trade-off between fraud stopped and customers interrupted, at any threshold
Alert queue โ†’ Payment Reasons, rule trace, similar past cases, the customer's message, and a case note in English or เฆฌเฆพเฆ‚เฆฒเฆพ (template by default; optional LLM draft that never decides and is rejected if it invents a number)
Cases One case per wallet with a review deadline (on time, due soon, overdue), a workload board per reviewer, filters, saved views and CSV export; a false-alarm verdict can say why
Customer appeals Customers who say a warned or held payment is genuine. A person answers every one; approving a hold releases it, and the verdict becomes a legitimate training label
Victim refunds A report on a completed payment opens a refund claim. After fraud is confirmed and two people freeze the mule, what is left in that wallet is paid back, shared by each victim's loss (PLATFORM.md ยง6)
Blocklist Wallets, phone numbers and domains known to be used for fraud. A listed wallet asks the sender to verify and wait; it never blocks money by itself
Webhooks ยท Partner API keys Signed, retried webhooks and customer SMS; scoped, rate-limited partner keys with a sandbox that returns every outcome on demand (PARTNER_API.md)
Follow a report (/track) A customer enters the reference they were given, gets a code on their phone, and sees where the report stands, in Bangla or English
Freeze approvals The second person of the two-person rule; approving a freeze can also settle open refund claims on that wallet
Network explorer ยท Mule rings ยท Agent risk Follow the money, see shared handsets, rank agents against their peers
Mule consortium Privacy-preserving cross-provider mule intel: OPRF tokens, signed daily bundles, partner lookups, disputes and a hash-chained audit log โ€” no MSISDN or handset leaves a provider (CONSORTIUM.md)
Model monitoring Performance by tier and scam type, drift, the model registry, shadow mode, the feedback loop
Fairness report ยท Decision policy ยท Audit log False alarms by group, the exact rules and texts in force, and who did what
Fraud types The eight kinds of fraud wallet customers in Bangladesh meet, what detects each and how well; check a suspicious message, verify a claimed payment against the ledger
Customer phone demo Plays the wallet app against the real scoring endpoint: send, warn, step-up, hold, report a scam, check a refund, and appeal โ€” side by side with what FraudLens decided
Guided tour A walk-through of the console in English or เฆฌเฆพเฆ‚เฆฒเฆพ (sidebar Take a tour), including API keys, blocklist, webhooks, appeals and refunds

๐Ÿ“Š Results

Every number comes from the 25-day test period (134,545 scored payments, เงณ35.5 lakh at risk), with thresholds fixed beforehand on validation data. Full tables are in the model card.

Each model against a rules-only system

A hand-written rule set, the kind many wallets run today, is right about one alert in eight. The served model is right about eight in ten at the same alert budget.

xychart-beta
    title "Ranking quality (PR-AUC) on the test period"
    x-axis ["Hand-written rules", "Anomaly only", "Mule model only", "Transaction model (served)", "Fusion of all three"]
    y-axis "PR-AUC" 0 --> 1
    bar [0.076, 0.175, 0.375, 0.834, 0.835]
Loading

The fusion is level with the single model here (0.8345 against 0.8343) and slightly behind it on validation and on precision, so the simpler model is served and the comparison is kept.

A scam type the models never saw

investment_scam does not exist in the training period and is built to defeat the easy signals: aged wallets, private handsets, money sent back to the victim to build trust.

xychart-beta
    title "Warn tier by scam type (bar = scams caught at the payment, line = money stopped incl. cash-out holds)"
    x-axis ["Account takeover", "Wrong send", "Lottery fee", "Impersonation", "Investment (unseen)"]
    y-axis "Percent" 0 --> 100
    bar [100, 100, 98.5, 96.8, 70.4]
    line [100, 100, 99.6, 100, 92.3]
Loading

Most of the unseen scam is still caught at the warning level, mainly through signals about the receiving wallet, and most of the rest is recovered at cash-out. At the stricter tiers it falls to 54.6% (step-up) and 39.5% (hold). The near-perfect bars on known types are a property of scripted synthetic fraud and should not be expected on real traffic.

How much money each tier stops

xychart-beta
    title "Victims' money stopped (bar = at the victim's payment, line = also counting held cash-outs)"
    x-axis ["Hold and above", "Step-up and above", "Warn and above"]
    y-axis "Percent of taka at risk" 50 --> 100
    bar [72.8, 78.1, 85.1]
    line [83.1, 90.1, 96.6]
Loading

The line assumes the hold on the mule's cash-out succeeds, so it is an upper bound on recovery, not a promise.

Where the interruptions land

98.5% of scored payments are never interrupted. Of the 2,012 that were:

pie showData
    title Interrupted payments in the test period
    "Hold (a person reviews)" : 1029
    "Warn (customer chooses)" : 661
    "Step-up (verify and wait)" : 322
Loading

Beyond single payments

Result
๐Ÿ•ต๏ธ Mule wallets 94 of 145 mules detected, 70 of them before any victim had paid, a median 40.2 hours ahead of the victim's report
๐Ÿ•ธ๏ธ Rings 34 rings covering 331 wallets; 94% of ring members are true fraud-cell wallets
๐Ÿช Agents No labels used: 90.9% precision in a review list of 22, finding all commission-farming agents
๐Ÿค Mule consortium At 25% shared-SIM (simulated), confirmed partner listings raise mule recall from 0.648 to 0.693 (~6 mules found only via partners) with almost no extra false alerts; เงณ paid to undetected mules falls by ~เงณ50k (CONSORTIUM.md)
๐Ÿ” Feedback loop Retraining on analyst verdicts raised PR-AUC from 0.774 to 0.912 on later data (reviewers were simulated and always right, so this is an upper bound)
๐Ÿ‘ฅ Shadow mode The retrained challenger scored all 136,180 served decisions without deciding any, and agreed on the tier for 98.5%; it would hold 1.06% of payments against 0.78%
๐Ÿ’ฌ Scam messages A classifier for English, Bangla and Banglish messages, with a link check, flags 90.2% of scam messages in new wording and 81.0% from scripts it never saw, at 0.23% and 0% of harmless ones. The corpus is synthetic; naming the exact fraud type on unseen scripts is weak (FRAUD_TAXONOMY.md)

๐Ÿ›ก๏ธ Trust and evaluation

A fraud system that cannot be checked is not one a bank can run. These are the checks in the repository.

The dashboards describe the service that is actually running. The whole test period was replayed through the live API and event stream, and make verify compared every served decision with the offline evaluation:

Check Result
Scored transactions compared 136,180 of 136,180
Tier differs 0
Feature rows that differ 0
Served in rules-only mode 0
Dead-lettered or failed events 0

It is fast enough to sit inside a payment request (one process, on a laptop):

Milliseconds p50 p95 p99
One decision in process: features, models, policy, reasons 3.1 3.4 n/a
Served, features + models + policy + reasons 1.9 2.3 3.6
Served, full round trip including the database commit 5.5 7.0 10.6

The first row is an average over a sample that is two-thirds alerts, which cost more to explain; the served rows are over real traffic, which is mostly allowed, and were measured with Postgres on the host. They move a lot with what else the laptop is doing: the same model measured 13.0 ms and 47.9 ms p95 on a busy machine (PLATFORM.md ยง11). The stream worker replayed 292,567 events in under six minutes, and the service is ready 8โ€“10 seconds after a restart with the feature state rebuilt exactly.

Guardrails that are enforced in code, not left to configuration:

Guardrail How it is enforced
The model never blocks money A policy whose hold tier lacks human review fails validation and will not load
A freeze needs two people Checked in the API and again by a database CHECK constraint
No training/serving skew One feature engine class builds the training table and serves live
No look-ahead Features use only state from before the transaction; a truncated-replay test checks it
The service answers without a model Rules-only fallback, which never holds a payment on rule points alone
Nobody edits history The audit log is append-only: a trigger refuses UPDATE, DELETE and TRUNCATE
Privacy by default Wallet numbers are masked (W***6128); revealing one writes the viewer's name to the audit log
The language model only words things Optional and off by default; its note is rejected if it contains a number or identifier that is not in the evidence

And the uncomfortable finding, stated plainly: wallets under 30 days old are interrupted on honest payments 6.0ร— as often as average when sending and 16.4ร— when receiving. A young receiving wallet is also the strongest honest sign of a mule, so the gap cannot simply be removed. The fairness report measures it so that a deployment can watch it.


๐Ÿ—๏ธ Architecture

flowchart TB
    subgraph OFF["Offline ยท make pipeline"]
        direction LR
        S["๐ŸŒ Simulator<br/>20,000 customers ยท 600 agents<br/>5 scam types, 1 held out"] --> FE["๐Ÿงฎ Feature engine<br/>61 features"]
        FE --> M["๐Ÿง  Models + registry<br/>LightGBM ยท Isolation Forest"]
        M --> P["๐Ÿ“œ Policy<br/>thresholds from alert budgets"]
    end
    subgraph ON["Online ยท make api"]
        direction LR
        IN1["POST /v1/score<br/>answer needed now"] --> SC
        IN2["POST /v1/events"] --> RS[("Redis stream")] --> SC
        SC["โšก Scorer<br/>same engine, same models, same policy"] --> PG[("PostgreSQL<br/>decisions ยท cases ยท audit log")]
        SC --> FEED["Live alert feed"]
    end
    P --> SC
    PG --> UI["๐Ÿ–ฅ๏ธ Console<br/>Next.js, no business logic"]
    FEED --> UI
    UI -->|"verdicts, freeze approvals"| PG
    PG -->|"verdicts as labels"| RT["๐Ÿ”„ Retrain โ†’ challenger<br/>shadow mode ยท drift"]
    RT -.->|"promoted only by a person"| M
Loading
What each layer does
Layer What it does
Simulator A 120-day world of wallets, agents and merchants with five fraud typologies, one of them held out of training, plus honest customers who look suspicious on purpose
Features 61 point-in-time features from one engine used for both training and serving
Models LightGBM transaction risk, mule-wallet score, Isolation Forest anomaly, agent peer comparison, ring detection
Decisions allow / warn / step-up / hold from versioned rules and budgeted thresholds, with a rules-only fallback
Explanations SHAP reasons in Bangla and English, rule trace, similar past cases, case note (template, or optional LLM draft fenced by grounding checks)
Platform FastAPI scoring API, Redis Streams ingestion, Postgres, roles, append-only audit log, appeals, refund claims, partner webhooks
Console Alert queue, cases, appeals, refunds, network explorer, mule consortium, ring freeze with second approval, agent risk, customer phone demo, guided tour
Consortium Simulated multi-provider mule intel over OPRF tokens and signed bundles; no raw identifiers leave a provider
MLOps Verdicts become labels, retraining, model registry, shadow mode, drift, fairness report
Tech stack and project layout
Layer Choice
Models Python 3.12, LightGBM, scikit-learn, SHAP
API FastAPI, SQLAlchemy, Alembic, PostgreSQL, Redis Streams and pub/sub, server-sent events
Console Next.js 16, React 19, Tailwind CSS v4, d3-force
Ops Docker Compose (Mac / Windows / Linux), Make or scripts/demo.*, uv, pnpm, Playwright, GitHub Actions
FraudLens/
โ”œโ”€โ”€ backend/
โ”‚   โ”œโ”€โ”€ src/fraudlens/
โ”‚   โ”‚   โ”œโ”€โ”€ simulator/   # the synthetic world and its fraud cells
โ”‚   โ”‚   โ”œโ”€โ”€ features/    # one feature engine for training and serving
โ”‚   โ”‚   โ”œโ”€โ”€ models/      # training, evaluation, registry, agents, rings
โ”‚   โ”‚   โ”œโ”€โ”€ decision/    # policy file, rules, tiers, reasons, case notes (+ optional LLM)
โ”‚   โ”‚   โ”œโ”€โ”€ consortium/  # privacy-preserving cross-provider mule intel (simulation)
โ”‚   โ”‚   โ”œโ”€โ”€ platform/    # scorer, cases, freezes, appeals, refunds, stream worker, audit
โ”‚   โ”‚   โ”œโ”€โ”€ mlops/       # verdicts as labels, retraining, shadow, drift
โ”‚   โ”‚   โ””โ”€โ”€ api/         # HTTP routes, schemas, middleware
โ”‚   โ””โ”€โ”€ tests/           # 435 tests
โ”œโ”€โ”€ frontend/            # the analyst console (incl. guided tour, phone demo)
โ”œโ”€โ”€ sdk/                 # partner API Python example client
โ”œโ”€โ”€ docs/                # architecture, model card, policy, platform, consortium, demo script
โ”œโ”€โ”€ docker-compose.yml
โ”œโ”€โ”€ Makefile
โ””โ”€โ”€ scripts/             # demo.sh / demo.ps1 (and up / demo-reset) for Mac, Windows, Linux

โšก Run it in one command

Requirements: Docker with Compose v2. Give Docker at least 4 GB of memory (Docker Desktop โ†’ Settings โ†’ Resources on Mac and Windows). Images are multi-arch (linux/amd64 and linux/arm64), including Apple Silicon.

Platform How to run the demo
macOS / Linux make demo or ./scripts/demo.sh
Windows (PowerShell) .\scripts\demo.ps1 โ€” Docker Desktop with the WSL2 backend
Windows (WSL2 / Git Bash) ./scripts/demo.sh or make demo
git clone https://github.com/0xdevabir/FraudLens.git && cd FraudLens
# macOS / Linux / WSL
./scripts/demo.sh          # or: make demo
# Windows PowerShell (Docker Desktop running)
.\scripts\demo.ps1

Then open http://localhost:3100 and sign in as analyst1, supervisor1 or admin. The password is generated on the first run and written to backend/.env as FRAUDLENS_SEED_PASSWORD.

The first start generates the dataset, trains the models, replays 25 days of traffic through the running service and retrains a challenger on the analysts' verdicts. That takes about ten minutes on a recent laptop, plus a few minutes to build the images, and progress is printed step by step. Later starts skip everything that already exists and are ready in under a minute.

Account Sees
analyst1, analyst2 alerts, cases, network, rings, agents, dashboards, the phone demo
supervisor1, supervisor2 the same, plus freeze approvals and the audit log
admin dashboards and the audit log, no customer data
# macOS / Linux / WSL
make down                  # stop, keep the data (Ctrl-C in the demo terminal also stops it)
./scripts/demo-reset.sh    # or: make demo-reset โ€” delete database, dataset and models
# Windows PowerShell
docker compose --profile demo down
.\scripts\demo-reset.ps1

The demo uses ports 3100 (console), 8010 (API), 5433 (Postgres) and 6380 (Redis), all bound to 127.0.0.1 (reachable as localhost on Mac, Windows, and Linux). Optional Compose variables are listed in .env.example.

Develop on the host

Needs uv, Node 22 with pnpm (corepack enable), and Docker for Postgres and Redis.

cp backend/.env.example backend/.env   # then set FRAUDLENS_SEED_PASSWORD (12+ characters)
make setup         # backend and console dependencies
make up            # Postgres and Redis
make pipeline      # data, features, models, policy, dashboard tables
make platform      # migrate, demo accounts, load the historical period
make api           # API and stream worker on http://127.0.0.1:8010

# in a second terminal
make replay        # test days 95-118 through the event stream
make replay-live   # day 119 one request at a time; writes the latency report
make verify        # served decisions against the offline evaluation
make review        # close the older cases with the simulation's ground truth
make retrain       # train a challenger on those verdicts (promotes nothing)
make console       # console on http://localhost:3100

To use a Postgres that is already installed on the host instead of the one in Docker, create a role with CREATEDB (the tests make and drop their own database) and a database it owns, set FRAUDLENS_DATABASE_URL in backend/.env, and run make redis in place of make up. backend/.env.example shows the URL and how to give Redis a password.

To run a challenger in shadow mode, set FRAUDLENS_SHADOW_MODEL_VERSION to its version (make models lists them) in backend/.env and restart the API. make help lists every target.

Check it
make test     # 230 backend tests; the platform tests need `make up`
make lint     # ruff, tsc, eslint
make smoke    # opens every console page as each role in a headless browser
make e2e-phone  # the customer phone in Bangla and English, plus an appeal round trip

make smoke and make e2e-phone need the demo (or make api and make console) running and make setup done. CI (.github/workflows/ci.yml) runs the tests and the lint against real Postgres and Redis, builds the console and builds both images.


๐ŸŽฌ Three-minute demo

Keep two browser windows open: one as analyst1, one private window as supervisor1. The full ten-minute walk-through is in docs/DEMO_SCRIPT.md.

Time Beat
0:00 Problem: the victim sends the money themselves, so the sender looks normal. Open the Executive summary: 25 days of traffic replayed through the live API, one scam type never shown to the models. (Optional: Take a tour in EN / เฆฌเฆพเฆ‚เฆฒเฆพ.)
0:20 Customer phone demo: an ordinary payment goes straight through. "Looks like a scam" shows the Bangla warning before the money moves. "Riskier" asks for a second check and a wait; skip the clock ahead to end it.
1:00 Hold: "Very likely fraud" is paused, not blocked. It appears at the top of the Alert queue over the live feed.
1:20 Payment page: the ranked reasons with the facts behind them, the rule trace, the closest past case. Switch the case note to เฆฌเฆพเฆ‚เฆฒเฆพ. Click a masked wallet number to reveal it.
1:50 Two-person freeze + refund: request a freeze as the analyst, approve it as the supervisor. Verdict confirmed fraud settles the victim's refund claim; on the phone, Check my refund shows the amount returned.
2:20 Mule rings ยท Consortium: open a ring of shared handsets, then Mule consortium to show partner listings as tokens (no raw numbers leave a provider).
2:40 Trust: the Fairness report shows young wallets pay more for false alarms; the Audit log shows the reveal from 1:20 and cannot be edited.
2:55 Close: the model ranks and explains, the customer gets the first say (warn, appeal, refund), and a person makes every irreversible decision.

๐Ÿ“š Documents

ARCHITECTURE.md How the parts fit and why they are split that way
DATA_ASSUMPTIONS.md What the synthetic world contains and what it leaves out
BANGLADESH_CONTEXT.md Sourced Bangladesh MFS evidence the simulator and pitch are grounded in
MODEL_CARD.md Models, results, ablations, fairness, limits
DECISION_POLICY.md Tiers, rules, thresholds, explanations, the language model's role
BUSINESS_CASE.md The threshold sweep as a monthly P&L in taka, assumptions and sensitivity
PLATFORM.md API, workflow (cases, appeals, refunds), security, stream, measured latency
CONSORTIUM.md Privacy-preserving cross-provider mule intel: OPRF tokens, bundles, disputes, measured lift
SCALING.md Several stream workers in one consumer group, measured; Kubernetes manifests in deploy/
FRAUD_TAXONOMY.md Eight kinds of fraud in the Bangladesh context, what detects each, the scam-message classifier and its limits
DEMO_SCRIPT.md The full walk-through (phone, freeze, refund, consortium, fairness)
INGEST.md The signed webhook for core banking: ISO 20022 pacs.008 mapping, HMAC signing, replay protection, decision callbacks
PARTNER_API.md Partner API keys, sandbox outcomes, signed callbacks; Python SDK under sdk/
SECURITY.md STRIDE threat model with the status of each mitigation, personal data and retention

โš ๏ธ What this is not

  • Everything runs on synthetic data, so absolute numbers will not transfer to real traffic. Fraud is more common and more scripted here than in real life, and thresholds must be re-fitted on real data. The mule consortium's four providers and shared-SIM rates are simulated assumptions; the crypto protocol is real code (CONSORTIUM.md).
  • It is a prototype, not a deployment. Stream workers scale out, but every one holds the whole feature state and events enter it one at a time, so throughput stops growing at a few workers (SCALING.md). TLS ends at a Caddy proxy, secrets come from files with no vault client, and keys rotate by command, not on a schedule (SECURITY.md).
  • The demo scaffolding is not the product. The demo accounts, the phone demo endpoints and the review simulator exist only outside production mode.
  • No model output moves or blocks money on its own. A hold waits for an analyst and a freeze needs two people. An optional language-model case note never decides; failed drafts fall back to the template (DECISION_POLICY.md ยง7).
  • The Bangla texts were written by the developers and have not been reviewed by a professional translator. The console has a sign-off per text for a translator to use; until it is filled in, they are shown as unreviewed.

The limits of each part are listed in its document under "Limits, stated plainly".

๐ŸŒ Impact

  • Customers: a warning in their own language at the one moment it can still help, a way to appeal a wrong interrupt, and a refund claim when money already left โ€” with far fewer honest payments interrupted than a rules-only system.
  • Fraud teams: one case per mule wallet instead of one alert per victim, each with its reasons and its network already laid out, plus partner mule signals that never expose raw identifiers.
  • The wallet operator: a measurable trade-off between money saved, customers interrupted and reviewer hours, and an audit trail for every decision.

Built by 0xdevabir

FraudLens: the victim presses send, so look at where the money is going.

About

In an authorised-push-payment scam, the victim presses "send" themselves. Their phone, PIN and location are all genuine, so checks on the sender see a normal payment. A caller pretending to be wallet staff, a "lottery fee", an "investment" with daily returns: the money goes to a mule wallet and is cashed out at an agent within minutes.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages