MENLO PARK, Calif.: WhatsApp begins testing Scam Alert, an optional feature that uses on-device machine learning to flag suspected scam messages.
The limited beta starts August 12. Meta is testing the feature with security researchers in WhatsApp’s Bug Bounty community. No timeline exists yet for a wider release, and initial regions remain undisclosed.
Once enabled, the feature downloads a machine learning model to the user’s phone. The model analyzes incoming messages from senders who are not in the user’s contacts. Classification happens entirely on the device. WhatsApp says “no message content leaves the device for classification,” and nothing is auto-reported to WhatsApp, Meta, or anyone else.
The model is trained on scam conversations previously reported by users. It checks whether messages match known scam patterns using linguistic signals and conversational structure. When it detects a likely scam, a private warning appears in the chat. Only the recipient sees it. The user can then block the sender, report the chat, or continue the conversation. Marking a chat as trusted removes the warning and prevents future flags for that conversation.
Users who believe a warning fired in error can optionally share their five most recent messages with WhatsApp to improve the system. That sharing requires explicit consent. Meta says it will also publish the model weights so independent researchers can verify the classifier is limited to scam detection.
The design addresses a long-standing tension for encrypted platforms. Regulators have argued that end-to-end encryption prevents safety protections. WhatsApp positions Scam Alert as evidence that fraud detection and encryption can coexist.
The feature follows device-linking scam warnings WhatsApp introduced in March. The company says the tool targets impersonation, social engineering, and AI-generated fraud lures.
