WhatsApp‘s new Scam Alert feature is now in limited beta, using an on-device machine learning model to warn users when incoming messages look like scam attempts. The privacy angle is the headline within the headline: none of the message content leaves the device, and the whole thing is optional.

The rollout is currently restricted to researchers in WhatsApp’s Bug Bounty community as the company tests the warning system. WhatsApp described the feature as ‘an on-device machine learning model to alert a user about potential scam messages,’ and was explicit that ‘no message content leaves the device for classification or is auto-reported to WhatsApp, Meta, or anyone else.’

How the WhatsApp Scam Alert Feature Actually Works

The model is trained on scam conversations that users have previously reported. When a message arrives from someone not in your contacts, it checks that message against known scam patterns using linguistic signals and probabilistic classification based on conversational structure. If it flags something, a chat warning appears giving the user three options: block, report, or continue the conversation.

Users who think a warning has misfired can mark the chat as trusted, at which point the warning disappears and Scam Alert will not flag that contact again. There is also an opt-in step at that point: if a user marks a chat as trusted, they can choose to share the last five messages received with WhatsApp to help sharpen the model’s accuracy going forward. Crucially, both the processed message data and the model itself stay on-device throughout, and the feature can be switched off at any time.

Meta put a finer point on the privacy architecture in an engineering team announcement on 12 August, as reported by Forbes: the feature ‘complements end-to-end encryption while enabling a user-controlled, optional scam alert when the model believes there’s a likely scam.’ That framing matters because it addresses an obvious tension: any system that analyses message content to catch scams sits uncomfortably close to a system that reads your messages. Running the model entirely on the device, and making it optional, is WhatsApp’s answer to that concern.

Part of a Broader Security Push

Scam Alert does not arrive in isolation. WhatsApp has been methodically adding security layers for some time, and this feature slots into a sequence of updates aimed at catching fraud before users engage with it.

In March, Meta announced that WhatsApp would warn users when behavioural signals suggest that a device-linking request may be fraudulent, a tactic commonly used to hijack accounts by tricking people into scanning a malicious QR code or sharing a linking code.

Two months before that, the platform began rolling out ‘Strict Account Settings,’ a lockdown-style security mode built for journalists, public figures, and other high-risk individuals facing sophisticated threats including spyware attacks. The feature came in direct response to a documented pattern: journalists, activists, and political figures had their phones infected with spyware, including NSO Group’s Pegasus, via WhatsApp in attacks exploiting zero-click vulnerabilities. Zero-click exploits allow attackers to compromise iOS and Android devices without any interaction from the target.

Taken together, the three features follow a loose hierarchy: Strict Account Settings hardens the account against targeted, high-sophistication attacks; the device-linking warning catches a specific account-takeover method; and Scam Alert covers the broad, high-volume end of the threat landscape, where automated scam messages hit ordinary users at scale.

WhatsApp is used by over 3 billion people across more than 180 countries, which makes it an obvious hunting ground for scammers. An on-device model that flags suspicious messages without sending content to any server is a reasonable engineering response to that scale, though how well it performs outside the current beta will only become clear once it reaches a wider audience. WhatsApp has said users can report incorrect flags and opt in to share messages to improve accuracy, so the model’s performance will, in part, be shaped by how people use those feedback mechanisms once the feature rolls out more broadly.

Share.

Software engineer and video game uber-nerd.

Comments are closed.

Exit mobile version