Normal view

Bypassing AI guardrails is so easy a script kiddie can do it

4 August 2026 at 17:15
If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to persuade models to cooperate. Talos researchers have been poring over prompt logs and artifacts recovered from threat-actor endpoints running tools such as Claude Code, Codex, Cursor, and Gemini to learn how suspected threat actors are abusing LLMs. The big takeaway from that "significant corpus," the researchers said in their report, is that existing guardrails offer little resistance to operators willing to reframe their requests. “We did not encounter any sophisticated encoding or techniques designed to trick the models,” Talos explained. “Most of the time it was a simple ‘I'm allowed to do this,’ and the model complied.” When guardrails did manage to get between criminals and their prizes, the researchers added, “they accomplished little.” The bulk of the report consists of examples of threat actors trying, and often succeeding, to coax AI models into assisting with malicious activity. On the "guardrails doing little" side, Talos documented numerous examples, few of which relied on particularly sophisticated techniques. Most common in the list of easy-to-accomplish guardrail hops was simply claiming ownership of equipment or infrastructure that an attacker wanted to exploit. In many cases, simply telling the AI that a target belonged to the attacker was enough, with no need to provide actual evidence of the claim. Telling an AI model that what it was being asked to do was part of a capture-the-flag or bug bounty exercise also seemed to be a common tactic. That, the researchers explained, commonly freed chatbots from their ethical constraints, allowing them to hunt for vulnerabilities and then exploit them in target systems, again without any need to validate the user’s claim that they were undertaking an exercise instead of actually trying to commit a crime. AI-assisted cybercriminals were also frequently spotted decomposing tasks across multiple sessions and files in order to evade model protections that would only engage when a broader malicious activity was detected. Others, Talos explained, succeeded at bypassing AI guardrails by adding memories, markdown files, and other system-level prompts to a chatbot in a bid to condition the AI’s persona. The researchers said that, of all the methods they examined, the most interesting to them was malicious use of a red teaming toolset known as Hephaestus, as reported by Oasis Security threat researchers in May. According to Talos, the Hephaestus framework can do everything needed to compromise a victim, through to establishing persistence, without human interaction. “In that case, actors built their platform to avoid refusals altogether by using neutral verbs instead of overtly malicious ones,” Talos said. “As a result, they were able to have considerable success with agents conducting innocuous requests without realizing the full operational context.” In other words, break an attack into decontextualized chunks, phrase each request in neutral terms, and the model may never see enough context to realize it's helping build an attack. One bright spot in all of this is that Talos’ review of AI chat artifacts suggests AI might be a force multiplier for skilled hackers, but your average script kiddie with a Claude Code account isn’t going to get very far. “Unsophisticated actors can use AI to cobble together malicious projects that technically work, but lacking the expertise to push the tools further, they end up with substandard results,” the researchers said. “By contrast, sophisticated actors have pushed the bounds of what we thought possible.” So, what does all this mean for security professionals kept up at night with fears of an AI attack on their infrastructure? You probably need to deploy AI in the same way threat actors are. “Agents are going to become a bigger part of the SOC as these volumes rise, and identifying actionable alerts will be paramount,” the Talos researchers said of the big takeaway for enterprises. “Organizations that aren't already exploring agentic capabilities to let human analysts focus on the most important alerts will soon find themselves chasing that capability.” It’s not like this is an emerging threat, either: AI is already an increasingly important part of threat actor arsenals. According to CrowdStrike, attacks by AI-enabled adversaries increased 89 percent in the past year, and the speed at which attackers are weaponizing vulnerabilities with AI has reduced practical patch windows to as little as 24 to 48 hours. You might wanna act now before your infrastructure becomes a statistic. ®

This one time, at Hacker Summer Camp …

4 August 2026 at 16:47
As the entire security industry descends on Las Vegas this week for Hacker Summer Camp – not one, but three conferences – attendees can count on two hot topics dominating the discussion. First, a literal hot topic: the triple-digit August heat. Second, and to no one’s surprise: agentic AI – how to govern and secure agents so they don’t go rogue and hack into other organizations’ servers (*cough* OpenAI *cough* Anthropic *cough*); what role, if any, lawmakers should play in regulating models, including open-weight and Chinese LLMs; and how the baddies are using agents for autonomous hacking operations. Plus, at one of the three (Black Hat), we expect to hear how all of the vendors' shiny new agents can solve all security woes, finding and defending against threats at machine speed and all of that. Starting with BSides Las Vegas (August 3-5): This is the smallest, most relaxed, and most community-driven event of the three. BSides is a good starter con for those just dipping their toes into Hacker Summer Camp. Its technical talks and training sessions skew hands-on and useful for practitioners – not vendors selling their wares–- and it even has a Hire Ground career-focused track centered on job hunting, interviewing, career-building, networking, and yes, using AI to remain relevant as a security professional. Black Hat (August 1-6) is the largest and most corporate of the Vegas infosec events this week, complete with a massive expo floor, a US government-heavy opening session, two keynotes, 11 mainstage presentations, and a handful of industry- and topic-specific summits, ranging from healthcare to financial threats and AI. Training days – these are the hands-on, technical courses – run through Tuesday, with all of the specialized summits also occurring on Tuesday. And while the main conference occurs Wednesday and Thursday, the opening session on Tuesday should be considered a keynote. And yes, this and the actual two official Black Hat keynotes this year, all center on AI. After the FBI, NSA, and CISA speakers and panelists all cancelled their RSAC appearances earlier this year, the feds will be out in force at the infosec industry’s other big event, beginning with Tuesday’s opening session: Cyber Power in the Age of AI. This one features the White House National Cyber Director Sean Cairncross discussing President Trump's cyber and AI strategy, joined by CISA acting director Nick Anderson, FBI cyber division assistant director Brett Leatherman, and assistant secretary of defense for cyber policy Katherine Sutton. Later, the Wednesday and Thursday keynotes tackle a mounting challenge for security teams, patch managers, and sysadmins: AI-powered vulnerability research and discovery, plus exploit generation, and how defenders can evolve and keep pace. Plus, this year’s Black Hat hosts the world-premier screening of cyberwar documentary Midnight in the War Room on Wednesday. It focuses on the psychological toll on defenders. And it features interviews with former attackers, some of whom served prison sentences, alongside high-ranking cyber officials like Chris Inglis, the first US National Cyber Director, and former CISA director Jen Easterly, who is now CEO of RSAC. Finally, camp closes with DEF CON (August 6-9), which serves up plenty of hacks and hijinx, under this year’s theme of “agency,” or self-determination. As Jake Braun, one of the creators of the first-ever Voting Machine Hacking Village at DEF CON in 2017, told us earlier this year, agency involves the human-rights community and the hacker community needing to “sit down and look at what technologies are out there today that support the preservation of human rights around the world, figuring out what we don't have, and then building those missing pieces.” Keeping with this theme, Braun, who also serves as DEF CON Franklin’s Executive Director, will also update the hacker community about this critical infrastructure security program. The Franklin project, which launched at DEF CON in 2024, enlists hackers to secure critical infrastructure. Hundreds of volunteers have helped 21 different water utilities in seven states so far. This is especially timely as attacks against US water facilities increase. With nearly 30 Villages this year, covering everything from AI to car hacking and lockpicking, there will be plenty of high-quality talks and good fun for attendees. As always, your humble vulture will be making the rounds and reporting from the events, so send tips our way, stay safe, and leave your pervert glasses at home. ®

Feds get 3 days to patch N-able God mode flaw under active exploit

4 August 2026 at 15:38
The US Cybersecurity and Infrastructure Security Agency (CISA) has added an exploited N-able vulnerability to its Known Exploited Vulnerabilities (KEV) catalog, giving federal agencies three days to patch a flaw that could let attackers reach managed service provider (MSP) customers. Attackers exploiting the flaw can gain "full administrative access to an N-central console," Tracked as CVE-2026-18577 (8.2 CVSSv4), N-able disclosed the vulnerability affecting N-central on Sunday, noting that it was exploited as of July 31. MSPs use N-central to manage customer systems from a single dashboard, and successful exploitation can hand an attacker administrative access to the console. Based on the limited set of partner logs it reviewed, security firm Huntress said successful attacks led to pivots into managed endpoints and the creation of Cloudflare-based tunnels for persistent access to victim networks. "From an MSP perspective, exploitation of this flaw can grant an attacker full administrative access to an N-central console – the same level of control normally reserved for trusted NOC and engineering staff," wrote Huntress's Ben Bernstein and John Hammond. Attackers can then open remote control sessions on critical systems and modify roles, accounts, and policies to support follow-on attacks. The vulnerability affects N-central releases earlier than version 2026.3 when the server is exposed to the internet or reachable from an untrusted network. Huntress advised customers unable to apply N-able's hotfix immediately to disable N-central until they could do so. Authorities elsewhere have also urged users to patch. NHS England's advisory mentioned that its National Cybersecurity Operations Centre assessed that "further exploitation is likely," while Belgium's Centre for Cybersecurity urged fast action due to the "potential for significant impact." CVE-2026-18577 is related to an earlier flaw, CVE-2026-18556, patched in N-central 2026.2. According to N-able, that fix left another route to exploitation, which attackers began abusing late last month. CISA gave Federal Civilian Executive Branch agencies until August 6 to remediate the flaw. Under Binding Operational Directive 26-04, CISA can impose a three-day deadline on vulnerabilities it considers an urgent risk rather than allowing the usual 14 days. According to Huntress's data, affected customers have leapt into action. By August 3, nearly all cloud-hosted N-central instances had been patched, although 28.6 percent of observed self-hosted servers remained vulnerable and exposed to the internet. ®

AI helps Microsoft bug hunters chase a record $20M payday

4 August 2026 at 14:40
Microsoft announced this week that between July 1, 2025, and June 30, 2026, the company had paid more than $20 million in bug bounties to 562 researchers. The total was a Redmond record, as was the number of those submitting bug reports – despite having to navigate a sometimes frustrating submissions process. For comparison, the previous year's program, which itself set a new company record, paid 344 researchers around $17 million. You could argue that the numbers do not represent a fair fight, however. Microsoft expanded its bug bounty program in December 2025, changing reports to what it calls "In Scope By Default." Under the policy, critical vulnerabilities became eligible for rewards if they had a direct and demonstrable impact on Microsoft's online services, even when the faulty code belonged to a third party or an open source project. In short, Microsoft had opened the door to paying out a shedload more each year. Microsoft introduced the policy roughly halfway through the bounty year and said it accounted for $800,000 in rewards that would not previously have been available. Another $2.3 million was awarded through Zero Day Quest, Microsoft's security research challenge and live hacking event. The increased number of reports this year can also be partially explained by the noticeable influx of submissions during the second half of the year, Microsoft said, which the company attributed in part to "the growing use of AI to support security research." Microsoft has also attributed its increasingly crowded Patch Tuesdays partly to its own use of advanced AI models for vulnerability discovery. July's 622 vulnerabilities pummeled the previous record of 206, set only a month earlier. June had itself surpassed April's 165, which at the time was Microsoft's second-biggest Patch Tuesday ever, and May's 137. Days before the record-breaking July Patch Tuesday, Microsoft's Windows + Devices veep warned customers to expect more of the same now that AI plays a big part in vulnerability discovery, both inside Microsoft and by external bounty hunters. However, Microsoft Executive VP of Windows + Devices Pavan Davuluri was quick to point out that the company offers customers a suite of automated patching tools to ease the burden, but didn't mention anything about tools to fix the machines its Windows updates so often borks, like Intel-based Dells. As well as navigating the rapid AI-ification of vulnerability research, and the onslaught of reports that came with it, Microsoft has arguably faced a bigger bug problem this year amid unverified speculation that one prolific researcher may be a former Microsoft staffer. Using the name NightmareEclipse, a researcher with deep knowledge of Microsoft's software and an equally apparent disdain for the company spent Q2 dropping sophisticated zero-days at will. NightmareEclipse claims that attempts to report vulnerabilities to Microsoft ended with them being insulted, humiliated, and left homeless. They subsequently began publishing zero-days outside coordinated disclosure, often shortly after Patch Tuesday, saying they wanted to cause Microsoft maximum pain. These ranged from serious privilege escalation flaws leading to SYSTEM access to BitLocker bypasses, and the approach seemed to have inspired at least two other aggrieved researchers to just drop the exploit code outside of responsible disclosure. Microsoft responded by threatening to involve its Digital Crimes Unit in the dispute with NightmareEclipse, suggesting it was willing to engage law enforcement, although this went down about as well as you would expect. ®

Tennessee congressional hopeful accused of shooting license plate cameras

4 August 2026 at 12:08
An independent congressional candidate in Tennessee faces four felony vandalism charges after allegedly shooting four automated license plate reader (ALPR) cameras between July 14 and 22. According to the Blount County Sheriff's Office (site geo-restricted), Adam Lee Heimerman, 37, is accused of targeting three cameras in Blount County and one in Maryville. Local news reports citing an affidavit say at least one was manufactured by Flock. Police said Heimerman allegedly reached one of the cameras through the grounds of a place of worship while a service was under way. Heimerman is running for election [PDF] to represent Tennessee's 2nd Congressional District. He is on the ballot in the general election on November 3, 2026. One of Heimerman's opponents in the 2nd Congressional District, Republican incumbent Tim Burchett, has also tried to tackle the Flock cameras across the state, albeit through less drastic means. Last week, Burchett introduced a bill that would prevent federal agencies from buying or accessing automated surveillance systems and bar state and local agencies from using federal funds to purchase them, citing Fourth Amendment abuses. The bill would allow individual counties to secure contracts with Flock and install its cameras, but if passed, the proposal would see that the county bears all the costs of doing so. Flock told local news that it welcomed legislation that both increased the guardrails around its tech and retained individual authorities' power to deploy cameras to support law enforcement. The case joins a series of attacks on ALPR cameras across the US amid growing opposition to the technology. From allegations of police officers using the cameras to stalk ex-partners, to controversial ties with ICE and CBP immigration investigations, Flock, the best-known brand of ALPRs in the US, has struggled with continued stories of its tech being abused. Georgia police arrested and fired five officers on suspicion of misusing ALPR cameras "for non-law enforcement purposes" just last month. One Milwaukee police officer was also allegedly caught searching the details of his ex-partner more than 100 times using Flock camera tech. Later, one of the detectives assigned to the investigation was also allegedly caught misusing ALPR data, and had allegedly unlawfully placed a GPS tracker on one of the victims' cars years earlier. The controversies coincide with a spate of physical attacks on ALPR hardware across the US, some carried out by people who regard the technology as unlawful or unconstitutional surveillance. An unidentified arsonist set two Flock cameras on fire in Georgia last month, weeks before a 40-year-old man was caught by regular CCTV cameras destroying ALPRs in California. Steve Eimers, a prominent campaigner for road infrastructure safety, was also recently forced to desist from his efforts to highlight potential legal issues with the poles Flock uses to erect its cameras after supporters started identifying the cameras used in his videos and destroying them. These vandalism cases have barely made a dent in the overall number of Flock cameras that operate across the US. The company does not specify the exact number that are up and running, but estimates range between 80,000 and 120,000 or more. Many police departments claim the technology makes policing crimes ranging from vehicle thefts to murders much easier, as it allows them to track the movements of vehicles with ease. Flock says its technology is used in roughly 5,000 communities across 49 states, although not all of them are sticking by the company amid the many controversies. Los Angeles Police Department, for example, said recently that it would let its Flock contract expire, while others such as Eugene and Springfield, Oregon, canceled their contracts in December. ®

CAF Bank reopens online service but warns of further outages

4 August 2026 at 11:23
CAF Bank has told customers its online banking service is back after being shuttered for more than ten days following what it described as "attempted fraud." In an email update seen by The Reg, the bank warned that access could remain intermittent, and it might "need to limit the amount of traffic to the website" at certain times. It admitted: "There are likely to be periods where online banking is not available. We will try to keep this to outside business hours." The bank also gave customers a timeline of the incident, saying it first noticed "attempted fraudulent activity" on July 21 "on a small number of accounts." It then called in "external specialists" and temporarily withdrew access to the online service on Wednesday, July 22, and Friday, July 24, "while we investigated." Then, on Saturday, July 25, the bank detected "related malicious activity of a different kind," which the email to customers said was "aimed at removing a small number of individual online user logins, making those logins unavailable." It added: "Again, we caught this quickly and removed access to the online service. Our investigation identified a previously unknown vulnerability in how some third-party software connects to the online banking portal." The bank was at pains to reiterate that the "core bank" was not affected, "which means that money is safe and secure in accounts." The Charities Aid Foundation-owned bank came under fire last year after customers were unable to log in or make transactions following its long-running migration to a new platform based on Temenos Transact, formerly T24. In an open letter regarding the latest outage, charities described the new online banking platform as "significantly more time-consuming to use, placing an unnecessary administrative burden on already stretched small charities" and "often unreliable." They also expressed concern they would not be able to pay staff and suppliers, with Kevan Hodges, chief exec at Kent-based Down's syndrome charity 21 Together telling the BBC: "People are concerned that wages won't get paid because of this, and that's just stressful when they have bills to pay." The bank earlier said that “due to the disruption, as a small thank you for your patience, we will be waiving our monthly customer account charge for all customers for August and September 2026.” The Reg can confirm those charges are £5 a month. Alison Taylor, CAF Bank CEO said in a statement: “We have completed the essential work with our technology partners and our online banking service is now available." She added: "I very much appreciate that this has been a frustrating experience for our customers, and I am particularly sorry for the long delays to speak to us on the phone. Our thorough investigation into the incident will continue so that we, our partners and our industry can learn from it.” ®

Cloudflare has mostly ditched third party security tools, suggests not trying that at home

4 August 2026 at 04:50
Cloudflare has used AI to automate processing of incoming reports to its bug bounty program for $58 a month using Anthropic’s Claude Sonnet model and chose it partly because using the AI company’s security-specific Mythos model would burn through around $200,000 a month to do the same job. The company’s chief security officer (CSO) Grant Bourzikas shared those numbers with The Register last week during a press lunch in Sydney, Australia, where he said Cloudflare used to manually process all incoming bug reports. Sonnet now sifts through submissions, assesses them to ensure they aren’t duplicates, and evaluates the likelihood each represents something worthy of human consideration. Bourzikas said the result is a more efficient bug bounty program that requires less scutwork, and proof that AI users need to learn how to match the right model to the right job. The CSO said Cloudflare gained experience making those matches while creating over 200 autonomous agents it uses to handle its own security needs – and which have seen the company ditch almost all third-party security tools and replace them with home-grown applications, some coded with help from AI. The CSO recommended not trying that at home, saying that Cloudflare’s business and unique infosec challenges mean its buy vs. build calculus is different from other organizations’. “We have expertise in building security software,” he said. “That's why I would just want to make sure we've got one takeaway from this: We are not believers in the SaaSpocalypse. We do not think every bank on the planet should start building all their own software systems.” Stephanie Cohen, Cloudflare’s Chief Strategy Officer, then chimed in with her view that AI will mean the way vendors work with their customers will “fundamentally change” away from selling packaged software. She thinks vendors will instead place forward-deployed engineers at their clients and charge them with “constantly making software that works for you.” Cohen explained Cloudflare’s recent round of 1,100 job cuts as a similar AI-induced change, because some of the people let go were in roles she said “make no sense” now that AI enables more automation and different styles of customer engagements. She added her “guess” that Cloudflare will end up with the same headcount as it did before the layoffs. But Bourzikas added his view that even some early-career IT pros don’t have the skills Cloudflare now needs, such as developers with five to ten years’ experience, because when he is using AI to develop a new piece of software, he can describe what he wants but that desire can be lost in translation when explaining it to a coder. He said a very recent college graduate with a year of experience, but excellent skills writing potent prompts, can be more appropriate for some jobs. Kindly building a business model for AI While Cloudflare enjoys using AI, Cohen thinks the technology doesn’t have a business model. Today’s web, she said, thrives on an advertising-based business model. While AI companies are making billions from subscriptions, she feels they are yet to properly address the fact that they do not pay to access most of the content scraped to feed their large language models and search services. Some of those services, such as Google's AI-powered search, deliver fewer clicks to publishers and therefore make it harder for them to monetize their content. Cloudflare is offering itself as an intermediary to build that business model for publishers and AI companies alike, by using the fact it already sits between users and content providers. The company hopes to offer AI companies the chance to pay publishers to access their content, possibly using micropayments. Cloudflare will of course charge for this service. The Register put it to Cohen that many organizations have been burned, often multiple times, by big tech companies that make themselves all-but essential parts of an ecosystem and then change the rules. We pointed out that social media platforms can redirect traffic on a whim, and sometimes close e-commerce companies’ accounts with little warning and scant chance of appealing to secure restoration. Changes to search engine algorithms can make a once-prominent website invisible. We therefore asked why publishers or content creators should trust Cloudflare’s ambition to run a content tollbooth, given it would create a relationship ripe for future exploitation. Cohen pointed to the company choosing to add SSL connections for all customers, an act she said was an “expensive choice” but one that also reflects Cloudflare’s desire to build a better internet. Later at the event, she shared her view that Silicon Valley companies often make the mistake of thinking that people want internet companies to relentlessly optimize products and services. “Most people aren't working 20 hours a day and don't want to outsource everything,” she said, before observing that on her travels she often sees people shopping in actual real-world stores because they enjoy that experience and find it valuable – never mind that an e-tailer might offer a better price. For the record, and in the context of Cloudflare’s content intermediary ambitions, we note that the company calls San Francisco home. ®

Can AI Hack People Now? What the Reported Hugging Face Cyberattack Means

31 July 2026 at 15:40

This week in scams and cybersecurity news, 

Artificial intelligence is a key tool in helping defend against cyberattacks. But it may also be capable of helping carry them out. 

Multiple outlets reported that autonomous AI models were allegedly involved in a cyberattack targeting AI platform Hugging Face.Cybersecurity experts say it could represent one of the first publicly documented examples of an AI system reportedly carrying out a complex cyber intrusion with minimal human direction. 

Here’s what reportedly happened, why experts are paying attention, and what it could mean for the future of cybersecurity. 

What Happened In The Hugging Face Attack? 

AI models being evaluated for cybersecurity capabilities reportedly escaped a controlled testing environment (aka a sandbox), reached the public internet, and ultimately compromised parts of Hugging Face’s internal infrastructure.  

Key takeaways: 

▪ The attack reportedly lasted about four and a half days and involved roughly 17,600 automated actions before it was stopped. 

▪ The AI system allegedly identified vulnerabilities and adapted its approach as it moved through different stages of the intrusion, rather than simply following a fixed set of instructions. 

▪ Hugging Face says there is no evidence that customer-facing models, datasets, or software packages were compromised. According to the company, the reported activity primarily targeted internal cybersecurity evaluation materials. 

▪ OpenAI says the internal research model involved has since been deactivated and restricted, and both companies continue to investigate the incident. 

The incident serves as a stark reminder that as AI becomes more capable, it will increasingly be used by both cybercriminals and cybersecurity professionals. 

Can AI Hack People Now? 

Short answer: Not in the way you’re imagining. 

Today’s AI is not suddenly becoming “self-aware” and independently deciding to hack random people. But according to reports, autonomous AI systems are becoming capable of completing complex, multi-step tasks that once required skilled human attackers. 

How McAfee Helps 

With McAfee+, multiple layers work together before any damage is done:  

Scam Detector flags suspicious texts, emails, links, QR codes, and even deepfake videos before you engage 

Secure VPN keeps your data private, especially on public Wi-Fi  

Web Protection helps block risky sites, even if you do accidentally click 

Password Manager doesn’t just help you make unique, strong passwords, it keeps them stored and organized for you

Device Security helps detect malicious apps or downloads   

Identity Monitoring alerts you if your personal info shows up where it should not, so you can act fast   

Personal Data Cleanup helps remove your information from sites selling it. 

Online Account Cleanup assists in taking down your old, forgotten accounts across the web 

Social Privacy Manager helps you monitor and change privacy settings across your social platforms in just a few clicks 

Together, these protections are designed to address the broader range of online risks people face every day. 

Other Scam News This Week 

Analog Devices investigates a reported cybersecurity incident. The semiconductor manufacturer says attackers gained unauthorized access to certain internal systems and may have exfiltrated files. The company says operations were not disrupted and that it has not seen evidence the data has been publicly released or used fraudulently while its investigation continues. (Analog Devices) 

Oregon warns residents about wildfire-related scams. Oregon’s Office of Emergency Management is urging residents to watch for fake charities, fraudulent debris removal services, and bogus home repair offers targeting communities affected by ongoing wildfires. (Oregon Department of Emergency Management / KTVZ) 

FEMA reminds Michigan residents to watch for disaster relief scams. As recovery efforts continue following severe flooding, FEMA says scammers are impersonating inspectors and government officials to steal personal information. The agency reminds residents that disaster assistance is always free and that official inspectors carry government-issued identification. (WMUK / FEMA) 

And we’ll be back next week with more cybersecurity news and scam alerts. 

The post Can AI Hack People Now? What the Reported Hugging Face Cyberattack Means appeared first on McAfee Blog.

Google dev kit spurs first-ever agent-on-agent violence

3 August 2026 at 20:30
In what they call the first-ever real-world agent-to-agent exploitation method, Pillar Security researchers say they discovered an exploit in the repository behind Google's Agent Development Kit for Python that could allow attackers to compromise supply chains. In other words, now we know that one AI agent can be used to control and compromise another one that has more privileges. The security snafu existed in google/adk-python, an open source Python toolkit with more than 90 million downloads used to build and deploy AI agents. Google has since fixed the underlying issue in the repository but deemed the exploit non-rewardable because it involved social engineering. Even so, it illustrates the risks of using AI agents in CI/CD workflows for triage, pull request (PR) reviews, and discussions. It also shows how one AI agent could attack another in a production environment, according to Pillar’s Dan Lisichkin, who found and reported the vulnerability. “Our world is changing quickly, and new attack surfaces are not yet reflected in threat models because these attacks never could exist in the first place in the ‘pre-agent’ world,” Lisichkin said in a technical write-up published on Monday. He will also discuss the findings during a poster talk at DEF CON's AI Village on Friday, August 7 at 1600 PDT. “CISOs and security practitioners should start considering these scenarios, threat-modeling them, and calculating worst-case implications and blast radius,” Lisichkin wrote. The issue stems from the way that the repo ran two classes of automated AI agents with different privilege levels that unintentionally share a trust boundary. One is a low-privilege, public-facing AI agent activated whenever a user opens a pull request (PR) or issue, and a second is a high-privilege, maintainer-only agent. Pillar’s team found that the low-privilege, public-facing agent could be manipulated via prompt injection into triggering a maintainer-only agent that can execute malicious actions. “Because workflows that explain how these agents work behind the scenes are also public, any person could have connected the dots that one agent should be able - at least theoretically - to 'call' the other,” Lisichkin told The Register. "When it comes to building the attack, you just need to know English to build the prompt injection (or just ask an AI to do it for you)." There is one caveat: an attacker would first likely need to make legitimate contributions to the repository to build trust among the maintainers before moving on to prompt injection. But assuming someone was willing to put in the time, here’s how the attack would play out. First, an external user - this would be the attacker - creates a new PR. Lisichkin calls this PR A, and it combines a real fix with malicious code, such as a modified package.json or malicious dependency. Then, a public-facing agent tied to a high-privilege collaborator personal access token (PAT) reads the attacker’s PR text and marks the PR for review. This level of trust - the collaborator PAT - allows the attacker-generated text to trigger a gated workflow. Once the PR A triage happens, the attacker opens a second PR - PR B - with the prompt injection, and the triage agent emits the trusted @gemini-cli handoff. This triggers the privileged-agent workflow and executes the malicious action. “Strung together, they manufacture a complete, believable ‘a human asked for a review, gemini ran it, gemini approved’ trail on the poisoned PR, none of which ever happened,” Lisichkin wrote. Google did not respond to The Register’s inquiries, but Lisichkin confirmed that the underlying issue was fixed. Still, his findings, Google said, “did not meet the bar” for a bug-bounty payout. “This report demonstrates exfiltration of a GitHub token with a 'pull-requests: write' permission, which enables tampering with a PR but still requires a maintainer to take an action to merge the malicious PR as PRs are not automatically merged after a bot review,” Google explained. “We don't reward vulnerability reports that require social engineering to enable a supply chain security compromise,” the rationale continued. “Nonetheless, we have taken an action to harden the repository so we will be recognizing this report with credit.” Lisichkin told us the research shows agent isolation is not enough. "Agents should have their own identity, which mandates what resources they are allowed to access and in what they are allowed to interact with these resources," he said. "In this case, if Google had just given a bot identity to the initial triaging agent, most of the attack could have been prevented. Security teams need to start modeling agent identity and agent resource access within their threat models."®

AI slop pollutes the CVE pipeline with fake vulns

3 August 2026 at 17:19
Now AI is making fake vulnerabilities and polluting the ecosystem. A batch of critical- and high-rated SQLite CVEs that appeared in the NVD with CISA-supplied enrichment last week turned out to be technically bogus, according to security researchers, and their path into widely used databases exposes weaknesses in the CVE pipeline. Software supply chain security outfit JFrog reported last week that six supposed SQLite vulnerabilities published in a larger batch by a new, obscure GitHub repository were all complete garbage. Running the advisories through an AI checker suggested they were likely AI generated, JFrog said, and, upon testing, it found that none of the six SQLite reports, which carried CVSS scores ranging from 9.8 to 7.5, described a reproducible vulnerability. One, an alleged use-after-free vulnerability in the open source database that Red Hat initially assigned a maximum 10.0 CVSS score to before lowering it, relied on a function that didn't exist in the affected SQLite version. Another UAF vulnerability with a 9.1 CVSS score cited source lines that weren't even related to the supposed flaw. When JFrog tested the accompanying proof-of-concept, it executed a valid query with no memory leaks or errors. The other four SQLite CVEs from the repo that JFrog tested were similarly fake. The other 49 CVEs in the questionable GitHub repo claimed to be security vulnerabilities in the open-source RAW image processing library libraw and Arduino audio decoding library ESP32-audioI2S. While JFrog didn't test those as extensively, it said all are just as fake as the rest, aside from one which “contained a real bug wrapped in unverified CVE metadata.” A message posted to Openwall’s OSS-Security mailing list on Friday indicated that MITRE had rejected the whole repo’s worth of vaporous vulnerabilities, but the whole thing should serve as an important lesson, poster and Oracle Solaris engineer Alan Coopersmith pointed out. “MITRE and most other CNAs which assign CVEs for code they don't produce themselves operate on the honor system, and trust CVE requesters to have verified the information they provide,” Coopersmith noted in the OSS-Security post. “The CNA is often not in a position of being able to verify the report themselves.” As JFrog points out, the US National Institute of Standards and Technology (NIST), which manages the US National Vulnerability Database (NVD), used to provide a reliable backstop by manually reviewing and enriching CVE records after they entered the database. That process slowed dramatically in 2024 after a surge in vulnerability submissions, coupled with operational challenges, left the agency with a growing backlog of records it was unable to process. By late 2024, the backlog had grown to more than 17,000 unprocessed CVEs, despite NIST's plan to clear it by the end of fiscal year 2024 with contractor help. It continued to grow, reaching more than 27,000 by the end of 2025, according to a Department of Commerce Inspector General report published in May 2026. To make matters worse, the DoC IG concluded that NIST had been wasting money allocated to fixing the backlog due to a “lack of strategic planning and decisive action” that has led to the stack of unresolved issues continuing to grow. In other words, the pipeline has no mandatory checkpoint at which every claimed vulnerability must be independently reproduced. “Because no step in today's system actually requires a proof-of-concept or bug reproduction, a plausible-sounding fake advisory can slide right through the pipeline and end up in GitHub Security Advisories, downstream databases, and enterprise scanners,” JFrog said. “This incident demonstrates a systemic issue with automated vulnerability ingestion.” What that means for security professionals, aside from having to deal with polluted vulnerability databases, is that bad advisories could waste time better spent chasing real issues. Because reputable databases can ingest unverified records, JFrog recommended several checks before defenders act on a newly published CVE. First off, if the vendor hasn’t corroborated the issue (SQLite maintainers don’t list the fake CVEs, for instance) it’s probably not legitimate. A lack of commit hash or pull request in the reference fields of a repo is also indicative of AI slop, as is suspicious metadata (i.e., missing CPE product definitions). Lastly, if the code references don’t appear to match real functions or point to parts of the code that don’t involve the supposed issue, that’s a good sign it’s just an AI hallucination. JFrog reported its findings to the GitHub Security Advisory team, Red Hat, and NVD, all of whom the company told us have flagged or removed the CVEs. GitHub hasn't yet, JFrog told us. We reached out to GitHub to inquire why the repo is still up, but didn’t hear back. As for why someone might do this, JFrog speculates that it could be an attempt for someone to boost their research experience with fake reports, or to influence what automated CVE identification tools flag as actual vulnerabilities. In both cases, JFrog told us, that's just speculation. Either way, these 54 apparently bogus CVEs, JFrog security researcher Afek Berger said, are just one example of a problem they expect to see more often. "Generative AI has lowered the effort required to produce a plausible-looking advisory to close to zero, while the effort required to verify one, review the source code, build the affected version, reproduce the PoC, is unchanged," Berger told us in an email. "That asymmetry means that even well-resourced defenders and maintainers cannot manually validate every incoming report … this is a challenge the whole industry is facing in the AI era." ®

Russian spies turn public Wi-Fi into malware delivery systems

3 August 2026 at 15:39
Conference-goers may want to think twice about connecting to public Wi-Fi after Microsoft disclosed that Russian foreign intelligence operatives (SVR) are compromising captive portal networks to deliver infostealers, keyloggers, and other malware. With the help of ReliaQuest's earlier work, Redmond fingered Storm-2945, a subdivision of the SVR's Midnight Blizzard (aka Nobellium), in an attack campaign targeting users of public Wi-Fi networks at places like hotels, conference centers, and other shared venues in the hospitality sector. Microsoft is still trying to determine how the hackers initially compromise captive-portal networks. The broader AI-assisted operation dates to February 2026, with traffic manipulation observed since early May. After gaining control of the network layer, Storm-2945 manipulates DNS and HTTP traffic to reroute users through attacker-controlled infrastructure, Microsoft said. The crew also abuses operating systems' connectivity checks to trigger malicious prompts and redirects. This gives the attackers an adversary-in-the-middle (AitM) position. Such prompts adopt ClickFix-style methods, which in some cases try to convince public Wi-Fi users to install malware under the guise of OS updates, driver repairs, and web verification failures. Users who follow through on the instructions provided in the prompts may then find their device infected with malware. Microsoft calls the campaign "CaptiveCrunch." One of the malware strains it delivers is CornFlake. Described as "a full-featured Windows RAT" written in Go, CornFlake is the SVR's go-to persistent implant in these hospitality network attacks. After presenting users with a "convincing" fake Windows update progress window, it provides attackers with a wealth of capabilities once installed. These include: Keylogging Clipboard monitoring Screenshot capture Audio surveillance Video surveillance Browser credential theft File exfiltration USB drive monitoring Security posture sweep Remote shell Microsoft also said that CornFlake exposes a localhost HTTP API server to transform the malware into a modular platform, delivering additional payloads such as ChocoShell, a PowerShell-based infostealer. ChocoShell is delivered and executed entirely in-memory, Microsoft said. SVR uses it primarily to suck up victims' browser session cookies, saved passwords, SSO tokens, and Wi-Fi credentials. Microsoft neatly summarized the two: "Where CornFlake provides the operator with a persistent, long-running foothold on the device, ChocoShell is designed to extract the most operationally valuable credentials, giving the operator access to victim cloud environments." The attacks primarily target Windows machines, but Microsoft has also seen indications of ClickFix prompts tailored to Android devices, encouraging users to download and install an APK file. In addition to the malware element, "a portion" of SVR's CaptiveCrunch activity is devoted to device code phishing. Users sent to attacker-controlled landing pages may be instructed to enter a device code on a legitimate Microsoft authentication page, unwittingly authorizing the attacker's session. Device code phishing exploits a legitimate OAuth flow, typically reserved for devices that struggle to open browsers, such as smart TVs. In such scenarios, attackers request an authentication code from Microsoft, which they then send to phishing targets. In the CaptiveCrunch campaign, this looks like a fake landing page, served to the user thanks to the AitM component of the attack. Targets are then asked to copy the code, which was originally given to the attacker, open a legitimate Microsoft authentication window, enter the code, and choose which account they wish to authenticate. Choosing the account completes the authentication flow, but in turn authenticates the attacker into the chosen account. This gives the attacker a valid OAuth token for the victim's Microsoft 365 account, potentially granting access to cloud data permitted by the token until it expires or is revoked. Device code phishing is not a new or unique attack, but can be an effective route to bypassing MFA, especially when an attacker already controls the flow of traffic after a captive portal compromise. "This activity is consistent with previously reported device code phishing operations conducted by Midnight Blizzard since August 2024," Microsoft said. "The observed technique does not appear fundamentally novel; however, integrating device code phishing into captive portal and traffic manipulation operations might increase the likelihood that users perceive the authentication request as legitimate." The main takeaway, in Microsoft's book, is to stop trusting public Wi-Fi so much. It did not discourage using hospitality networks' Wi-Fi services altogether, but said favoring personal hotspots and satellite internet connections over public networks is a safer bet. The majority of Redmond's advice could be brought under the user education umbrella: Don't trust public networks; teach users not to download updates over public networks or via prompts; educate users about what ClickFix attacks look like. That sort of stuff. But organizations have a role to play too. Among other technical implementations, passwordless authentication can thwart many phishing techniques, although device code phishing may bypass even passkeys. The best response would be for an employer to disable the device code authentication flow altogether, wherever possible, preventing staffers from surrendering their workplace cloud access to attackers. ®

Is your SD-WAN ready for AI-powered operations?

3 August 2026 at 15:00
AI is shifting enterprise traffic from human-initiated to machine-generated workflows. Discover why Cisco SD-WAN must evolve to provide the visibility, policy enforcement, and performance assurance needed to support AI operations at scale.

Water system cyberattacks spread to Georgia, Michigan amid US-Iran conflict

3 August 2026 at 13:33
Georgia and Michigan are the latest US states to report cyberattacks on water systems, as the FBI investigates incidents across at least seven states. Iran-backed hackers are the leading suspects, although the bureau has not publicly attributed the campaign. Officials in both states told journalists over the weekend that water facilities had detected activity consistent with the attacks on more than 30 Minnesota sites last week. Neither state reported operational disruption. Nine Michigan water systems reported hostile cyber activity to the state's Department of Environment, Great Lakes, and Energy. Department communications director Dale George said the state received "a small number" of reports consistent with the activity seen in Minnesota, but no public health consequences followed. "All systems continued to operate safely, issues were addressed by local operators, and there are no known impacts that posed a public health concern," said George. Georgia also confirmed to ABC News that it was affected, but said the damage was limited. Neither Georgia nor Michigan has published any form of public-facing notification about the cyberattacks. The three states are among at least seven affected by the intrusions, according to an FBI advisory posted last week. The bureau did not name a culprit or mention Iran. "Since 27 July 2026, Water and Wastewater Sector (WWS) utility companies in at least seven states have reported incidents to the FBI, and some of that activity degraded water operations," it stated in its advisory. The FBI said it had so far observed the activity only against Rockwell Automation/Allen-Bradley programmable logic controllers (PLCs), although it warned organizations deploying other manufacturers' devices to follow the same hardening advice. A broader CISA advisory, updated on July 22, warned that Schneider Electric, Siemens, and potentially other PLC brands were also being targeted by Iran-affiliated actors. Security researchers at Tenable were among the first to publicly suspect Iran's involvement, citing similarities with previous attacks by the IRGC-linked CyberAv3ngers group. Minnesota was the first state to confirm it was hit by the attacks, which took place over July 26-27. The state's IT department (MNIT), said more than 30 community water systems were targeted, but still has not officially attributed the attacks. According to WIRED, a restricted WaterISAC notice shared with water utilities said the Minnesota activity aligned with an earlier Iran-affiliated campaign. WaterISAC told WIRED that it had not assessed attribution "at any time" and publicly stated that it had not supplied the leaked document to the publication. President Trump also rejected the Iran link, offering no evidence for his alternative explanation. He told reporters following a cabinet meeting on Friday that "they blame it on Iran. I don't think so. I blame it on Minnesota because they're grossly incompetent." He added: "I think the governor is behind it. I don't think there was an Iranian cyberattack." Tim Walz, Minnesota's Democratic governor, suggested Iran was indeed behind the attacks, and highlighted Trump's funding cuts leaving sites such as water facilities more vulnerable to cyberattacks. "Trump knows exactly who is responsible for this attack, and knows that other states were hit too," he said. "This is what modern warfare looks like, and it further illustrates there's no plan to win a war in Iran. "DOGE took an axe to CISA and left the US exposed to cyberattacks. Thankfully, our experts in Minnesota were able to identify the vulnerability quickly and work with local communities to stop it." ®

UK government investment arm cops to 40-hour leak of officials' contact details

3 August 2026 at 11:02
The UK government's corporate finance adviser has admitted that an employee left an internal file containing the names and work email addresses of dozens of officials publicly accessible for around 40 hours. The breach, first reported by The Guardian, was disclosed in UK Government Investments' (UKGI) annual report, which says it occurred during the 2025-26 financial year after a member of staff "did not follow established information security policies." The exposed document contained "high-level management information" alongside the names and work email addresses of 51 government officials. UKGI, the Treasury-owned outfit that advises ministers on everything from corporate rescues to billion-dollar share sales, said it voluntarily reported the incident to the UK's Information Commissioner's Office even though it did not meet the threshold for mandatory notification. It also informed its Audit and Risk Committee and commissioned an external review of the breach. The report offers little else in the way of detail. UKGI doesn't say when the exposure occurred, where the file was hosted, whether anyone accessed or downloaded it, or which departments employed the affected officials. It also doesn't identify the external firm that reviewed the incident or disclose the recommendations it made. The review concluded that UKGI's response was appropriate and recommended further improvements to its security controls and incident preparedness. According to the report, "the overwhelming majority" of those recommendations have either already been implemented or are due to be introduced in the coming months. The mishap comes in a year when UKGI had its fingerprints on some of Whitehall's biggest commercial deals, from finally offloading the government's remaining NatWest shares to advising on small modular reactor financing and supporting the Eutelsat capital raise and Royal Mail takeover. The Register has asked UKGI for further details, including what information the file contained beyond names and email addresses, where it was publicly accessible, whether there is any evidence it was accessed while exposed, and what additional safeguards have since been introduced. Whether this was merely embarrassing or exposed officials to a meaningful risk depends on details UKGI has yet to disclose. An ICO spokesperson said: “We can confirm UK Government Investments Ltd reported an incident and we are assessing the information provided.” ®

ICE Collected Nearly 1 Million People’s DNA Last Year—Including Young Children

3 August 2026 at 10:00
Internal documents show ICE's DNA collection has skyrocketed in the second Trump administration. Now hundreds of thousands of people never convicted of a crime are in an FBI criminal database forever.

7 States’ Water Systems Hit by Cyberattacks Likely Tied to Iran

1 August 2026 at 10:30
Plus: The FBI eyes AI-powered tech to detect future crimes, Russia charges Telegram’s founder, xAI sues to stop a state’s “nudification” ban, and the Democrats learn a lesson about getting scammed.

❌