WhatsApp launches new scam alert feature to combat social engineering
Key Takeaways WhatsApp has introduced an optional “Scam Alert” feature to detect potential scam messages. The system uses an on-device machine learning model, preserving end-to-end...
Key Takeaways
- WhatsApp has introduced an optional “Scam Alert” feature to detect potential scam messages.
- The system uses an on-device machine learning model, preserving end-to-end encryption by processing message content locally.
- Warnings appear only to the recipient, who retains full control over reporting or blocking.
- A confidential federated analytics pipeline, built on Trusted Execution Environments (TEEs), aggregates anonymous performance data without compromising user privacy.
- Model versions are published to a third-party transparency ledger to prevent manipulation.
WhatsApp, the Meta-owned messaging giant, has unveiled a new, opt-in “Scam Alert” feature designed to shield users from increasingly sophisticated social engineering attacks. This innovative tool leverages on-device machine learning to identify potential scam messages while rigorously upholding the platform’s commitment to end-to-end encryption.
Table Of Content
On-Device Intelligence Safeguards User Privacy
As cybercriminals adapt their tactics from basic impersonation to advanced AI-generated lures, WhatsApp acknowledges the need for equally rapid evolution in its defensive capabilities. Scam Alert represents a significant stride in this ongoing effort.
Upon activation, the feature downloads a compact machine learning model directly to the user’s device. This model then autonomously analyzes incoming messages from non-contacts, scrutinizing conversational structures and linguistic patterns for indicators consistent with known scam methodologies. Crucially, all message content analysis occurs locally on the device, ensuring that no private communications are transmitted to WhatsApp, Meta, or any external entity for classification. User privacy is further protected by the system’s design, which mandates explicit user action for any message to be reported.
Should the model flag a message as a probable scam, a discrete warning appears within the chat interface, visible only to the recipient. This alert provides users with clear options: block the sender, report the message, continue the conversation, or dismiss the warning if they deem it a false positive.
Architectural Principles: Privacy, Control, and Transparency
The Scam Alert feature is founded on three core tenets: exclusive on-device processing, the absence of automatic reporting, and complete user autonomy. To evaluate the efficacy of Scam Alert without infringing on privacy, WhatsApp developed a confidential federated analytics pipeline. This system operates within Trusted Execution Environments (TEEs), specifically confidential virtual machines, to aggregate performance metrics.
This pipeline meticulously collects only anonymous statistical counts, such as the number of warnings issued and the subsequent user actions. Before any data reaches Meta’s servers, differential privacy noise is applied, further obscuring individual data points. The end-to-end process, from on-device data minimization through OHTTP relay job selection and RA-TLS attested orchestrator and aggregator TEEs, culminates in the output of differentially private, anonymized statistics.
Mitigating Model Manipulation and Ensuring Integrity
A critical security challenge for any system relying on server-delivered models is the potential for targeted manipulation by malicious actors or insiders. Meta addresses this by publishing every version of its machine learning model, identified by a unique SHA-256 hash, to a third-party, append-only transparency ledger prior to deployment.
Model download requests are routed through an OHTTP relay, which strips IP addresses and authenticates via anonymous credentials. This design prevents the server from linking model requests to specific users. Furthermore, the assignment of users to experiment groups for testing new model variants occurs entirely on-device, leveraging locally generated randomness to prevent the server from influencing which model a user receives.
WhatsApp’s threat model comprehensively considers external attackers, malicious insiders, and compromised supply-chain vendors. Defenses include TEE code isolation, encrypted DRAM, CVM hardening, and stringent restrictions that preclude even Meta engineers from gaining runtime shell access to the confidential computing environment.
User Auditability and Community Collaboration
Users can independently audit the system’s operation through an in-app transparency log. Accessible via “Account,” “Request Info,” and “Scam Alert Activity,” this log details which messages were scanned and the specific model version employed. WhatsApp is also expanding its Bug Bounty program to encompass the model weights and the federated analytics pipeline, actively inviting external researchers to validate that the system’s functionality is exclusively focused on scam detection.
Scam Alert is currently rolling out in a limited Beta phase. WhatsApp intends to rigorously stress-test the system alongside its security research community before a broader public release. The company also plans to release a comprehensive engineering white paper detailing the pipeline’s design, building upon its previously peer-reviewed PAPAYA Federated Analytics Stack work presented at USENIX NSDI 2025. This deliberate, phased introduction underscores an industry-wide commitment to pairing privacy-preserving AI features with independently verifiable transparency mechanisms, moving beyond reliance solely on internal assurances.
What You Should Do
- Enable Scam Alert: Once available, activate the Scam Alert feature in your WhatsApp settings to benefit from proactive scam detection.
- Exercise Caution: Always be skeptical of unsolicited messages, especially those from unknown contacts, even with Scam Alert enabled.
- Verify Warnings: Pay attention to Scam Alert warnings. If a message is flagged, carefully consider the sender and content before proceeding.
- Report Suspicious Activity: Utilize the in-app options to block and report suspicious messages to WhatsApp.
- Review Transparency Log: Periodically check your in-app transparency log under “Account” > “Request Info” > “Scam Alert Activity” to understand how the feature is operating.
- Stay Informed: Keep your WhatsApp application updated to ensure you have the latest security features and model versions.
Disclaimer: HackersRadar reports on cybersecurity threats and incidents for informational and awareness purposes only. We do not engage in hacking activities, data exfiltration, or the hosting or distribution of stolen or leaked information. All content is based on publicly available sources.



No Comment! Be the first one.