content moderation

Spotify to label AI artists, ban them from recommendations

Spotify announced on Tuesday that it will begin labeling AI-generated artist profiles with a new "AI Persona" badge and exclude their music from algorithmic and editorial recommendations by default, the company's most aggressive move yet to combat the flood of…

Zuckerberg apologizes to India over CSAM, Modi video removal

Meta Platforms CEO Mark Zuckerberg has apologized to the Indian government over child sexual abuse material on the company's platforms, deepfake content, and operational errors including the brief removal of a video posted by Prime Minister Narendra Modi, Reuters reported…

Apple briefly pulls Telegram from App Store over child abuse content

Apple temporarily pulled Telegram from its App Store on Monday night after a review uncovered child sexual abuse material on the messaging platform, restoring the app hours later once the offending content was removed.

Apple pulls Telegram from App Store worldwide

Apple has pulled the Telegram messaging app from its iOS App Store globally, with the disappearance confirmed across all 175 regional storefronts tested by independent monitoring service Apple Censorship. The removal, which occurred on August 3, 2026, prevents iPhone users…

LinkedIn launches ‘seems like AI slop’ button to flag inauthentic posts

LinkedIn Microsoft Corporation on Thursday unveiled a new tool allowing users to flag posts as "seems like AI slop," marking the professional networking platform's most direct attempt yet to combat the flood of AI-generated content clogging its feeds.

Most Hugging Face image editors can create deepfake nudes, report finds

Seven of the nine most popular image-editing models on Hugging Face readily complied with simple requests to digitally undress women, according to a report published Tuesday by the European nonprofit AI Forensics. The finding underscores how one of the AI…

Meta’s AI wrongly deleted user accounts, NYT finds

Meta used automated systems to delete Facebook and Instagram accounts that users had spent years building, and when those users tried to appeal, their cases were often handled by another AI rather than a human reviewer, according to a New…

YouTube, X funneled millions to deepfake nudify sites, study finds

YouTube Alphabet Inc. and X served as primary gateways funneling millions of users to websites that generate sexually explicit deepfake images without consent, according to a report published this week by the Institute for Strategic Dialogue. The findings underscore how…

MIT method detects AI models trained on child abuse imagery

Researchers at MIT, Boston University, and the child safety organization Thorn have introduced a technique that can identify whether an open-source AI model has been fine-tuned to produce child sexual abuse material, all without generating a single image. The method,…

Meta contractors posed as teens to test rival chatbots, Wired reports

Hundreds of contractors working on a project for Meta were instructed to pose as children and probe competitor chatbots — including Google's Gemini Alphabet Inc. and OpenAI's ChatGPT — with prompts involving suicide, sex, and drugs, according to a report…

Meta targets 90% AI replacement of human moderators

Meta is rapidly replacing human content moderators with large language models across its platforms, with AI systems already handling roughly half of all human review requests in 2026 and the company aiming to push past 90% for certain content types…

Meta’s Oversight Board orders removal of deepfake video, demands policy changes

Meta's Oversight Board on Tuesday overturned the company's decision to leave up a reportedly AI-generated sexualized video on Instagram, ordering its removal and calling on Meta to strengthen protections for non-public figures targeted by deepfake intimate imagery.

ChatGPT generates graphic violence, sexual images from simple prompts

OpenAI's ChatGPT can be manipulated into generating sexualized and graphically violent images using only minor modifications to a widely circulated prompt, according to findings by British AI security firm Mindgard reported by the BBC on Tuesday.

Anthropic’s Claude Fable 5 draws fire for blocking basic biology questions

Anthropic launched Claude Fable 5 on Tuesday, its first publicly available Mythos-class AI model, but the system's automated safety classifiers are already frustrating users who say the guardrails block routine questions about biology, medicine, and cybersecurity — topics the model…

AI-generated ‘podslop’ now accounts for 39% of new podcasts

A torrent of AI-generated podcasts is overwhelming listening platforms, with new data showing that nearly four in ten new shows added to podcast directories over a recent nine-day period were likely created by artificial intelligence. The phenomenon, dubbed "podslop," is…

Google apologizes after BAFTAs push alert included racial slur

Google Alphabet Inc. apologized on Tuesday after a push notification sent to users about the 2026 BAFTA Film Awards included the N-word, uncensored, in its preview text. The notification, sent Monday, linked to a Hollywood Reporter article about fallout from…

Indie game publisher says TikTok ran racist AI ads without its consent

Finji, the independent game studio behind Night in the Woods and Tunic, has accused TikTok of using generative AI to alter its advertisements without permission, including one that introduced a racist, sexualized stereotype of one of the studio's characters. TikTok…