WhatsApp has started a limited beta rollout of Scam Alert, an optional protection feature designed to spot suspicious messages coming from people you don’t have saved as contacts. The key idea is that the detection happens locally on your phone, rather than sending message content to WhatsApp or its parent company. For users, that means fewer privacy trade-offs while still adding an extra layer of caution.
Below is what’s known about how WhatsApp Scam Alert works, how warnings are handled in chat, and what safeguards the company says it uses to keep the system trustworthy during the beta period.
Scam Alert runs classification on your device
According to WhatsApp, the feature relies on an on-device machine learning model. After you enable the option, WhatsApp downloads the model to your device. From that point onward, incoming messages are evaluated for patterns associated with known scam behavior.
WhatsApp describes the approach as working alongside end-to-end encryption. In its design, the classification is performed entirely on the device, and there is no automatic reporting of message content back to WhatsApp or Meta.
Instead of sending the full text, the model looks at signals such as conversational structure and linguistic cues—signals that can help identify messages that resemble common scam patterns.
What happens when a message is flagged
If WhatsApp Scam Alert decides a message is suspicious, it shows a warning only to the recipient inside the chat. The sender does not receive any notification that their message was flagged.
From there, WhatsApp gives you several options, so you remain in control:
- Block the contact
- Report the message
- Ignore the warning
- Mark the conversation as trusted, which suppresses future alerts for that chat
This “trusted conversation” option is meant for situations where you still receive messages that the model might incorrectly interpret as risky, but where you know the context is legitimate.
How WhatsApp improves accuracy without sharing content
During the beta, WhatsApp needs a way to confirm the feature is working. It says it built a “confidential federated analytics pipeline” to measure performance while keeping message content on-device.
WhatsApp states the pipeline collects only two categories of data:
- How often warnings were triggered (counts)
- What action users took afterward (for example, blocking or marking a chat as trusted)
Because the system does not transmit message content, WhatsApp frames this as a privacy-preserving way to understand whether the warnings are useful and how users respond.
Transparency for model updates and anti-targeting protections
Machine learning features can become risky if an attacker is able to alter the model a specific person receives. WhatsApp says it has an additional process intended to prevent a scenario where one model version could be secretly pushed to a specific user.
Every model release, WhatsApp explains, must be logged on a third-party, append-only transparency ledger before it can be distributed. Along with the release, WhatsApp publishes a manifest containing SHA-256 hashes covering the model weights and related files.
The manifest digest is signed using Ed25519 keys held by Cloudflare—not by Meta. Before the model runs, devices verify the signature, cross-check the information against the ledger, and confirm the downloaded files match the published hashes.
In practical terms, this is meant to ensure the model package a device receives matches what was publicly recorded for that release.
Users can review Scam Alert activity
WhatsApp is also adding a way for users to examine what the feature did on their phone. Users will be able to access a transparency log inside the app.
The log is described as reachable through:
Account > Request Info > Scam Alert Activity
Within that area, WhatsApp says users can see which messages were scanned, the outcome, and which model version made the decision. This gives more visibility into how the feature behaves over time.
Opt-in beta design and continuing development
WhatsApp describes this rollout as an early technical preview rather than a finished product. That matters because it signals that the detection model and the overall behavior of Scam Alert may change based on feedback from researchers and users during the beta period.
Also, the feature is optional. Once enabled, the model remains on-device and can evaluate incoming messages under the rules described above.
Threat model and defenses against system-wide compromise
WhatsApp’s documentation includes a threat model covering multiple kinds of risk, including external attackers, compromised insiders within infrastructure, and supply-chain risk. The company argues that the pipeline’s defenses are intended to ensure that targeting a single user’s data would require compromising the entire system rather than just one endpoint.
While users may not interact with these controls directly, the company is effectively saying the privacy and integrity measures are designed to reduce the chance that the feature becomes a surveillance tool.
Bug bounty coverage expands
To support security scrutiny, WhatsApp says it is expanding its bug bounty program to include Scam Alert. That suggests the company expects external researchers to test for weaknesses related to the feature’s model delivery, on-device behavior, and related user flows.
Automatic key verification on Signal (related context)
In related news, Signal introduced automatic key verification. The goal is to complement its existing safety number system by enabling verification without requiring an in-person meeting or an extra communication channel.
Signal’s approach is designed to detect cases where a user’s public key is swapped without their knowledge—such as when an attacker compromises infrastructure and associates a different key with a target’s phone number. The feature is opt-out and includes a system called key transparency, with independent auditors supporting the log.
This is not part of WhatsApp Scam Alert, but it reflects a broader trend across encrypted messaging apps: adding user-friendly safety mechanisms while strengthening protections and auditability.
Conclusion: A privacy-forward warning system in beta
WhatsApp Scam Alert is an on-device feature aimed at flagging suspicious messages from non-contacts. WhatsApp’s messaging focuses on local classification, no automatic content reporting, and controls that let users review the feature’s activity.
With transparency ledger logging for model integrity, confidential analytics that collect only counts and user actions, and expanded bug bounty coverage, WhatsApp is positioning Scam Alert as a beta-first system built to evolve. If you enable it, you’ll get warnings inside your chats—while keeping the decision-making and message content largely on your own device.
Source: https://www.securityweek.com/whatsapp-unveils-new-scam-alert-feature/
