On-device Scam Alert model flags suspicious chats without sending message content to Meta
Meta has begun testing an AI-powered feature on WhatsApp designed to flag conversations that show signs of fraud before users get pulled further in.
The optional tool, called Scam Alert, is rolling out to a limited group of beta users, according to The Verge, and displays a warning banner directly inside a chat when the system detects patterns commonly associated with scams.
Scam Alert is different from other AI-powered moderation systems in that it uses its machine learning algorithm on the user's device and does not upload the content of messages to the server of Meta for analysis.
In an Wednesday's engineering announcement, Meta said the approach "complements end-to-end encryption while enabling a user-controlled, optional scam alert when the model believes there's a likely scam."
The content of messages is not uploaded anywhere for classification unless it is explicitly chosen by the user to do so.
Upon detecting a chat, a notification is raised which is viewable solely by the recipient of that message and not the other person involved in the chat. Thereafter, the user has options for blocking the contact, reporting the account, or proceeding with the normal course of action.
In the case where the notification is incorrectly raised, labelling that chat as trustworthy will remove the notification and prevent Scam Alert from raising that notification again. Also, those users that wish to aid in improving the model may choose to submit the last five messages from such a chat.
Scam Alert builds on a separate system WhatsApp introduced earlier this year to catch suspicious requests to link new devices to an account, a common tactic scammers use to hijack access without needing a password.
Meta has also said it removed more than 159 million scam-related ads last year, 92% of them before users had a chance to report them, alongside 10.9 million accounts across Facebook and Instagram tied to criminal scam operations.