On an afternoon in mid-May, a palpable sense of urgency permeated a virtual gathering and a conference room at Microsoft’s sprawling Redmond, Washington, headquarters. Dozens of the tech giant’s engineers and their managers convened to discuss "Project Glasswing," a critical initiative born from a burgeoning challenge: the ability of a new artificial intelligence model, known as Mythos, to uncover weaknesses in Microsoft’s software at a speed that outpaced the company’s ability to remediate them. Developed by the AI powerhouse Anthropic, Mythos had been granted access to select organizations creating software integral to global operations, from everyday consumer applications to critical government infrastructure. The objective was clear: to identify and fix these vulnerabilities before malicious actors, including nation-states like China, could leverage similar advanced AI tools for espionage and sabotage.
The meeting’s central question, posed by an engineer, cut to the heart of the matter: "Did Mythos live up to the hype that Anthropic claimed it would have had?" A manager’s affirmative response, captured in a recording reviewed by ProPublica, confirmed the AI’s remarkable efficacy. The version in use, Claude Mythos Preview, was identifying bugs at an "unprecedented clip," according to internal documents, pushing Microsoft engineers into a "mad dash" to close the widening gap between discovery and resolution.
The Mythos Deluge: A Cascade of Critical Flaws
The scale of the problem was starkly illustrated by a presentation slide detailing the month of April alone. Mythos had unearthed 90 "critical" bugs and 141 "important" ones within SharePoint, Microsoft’s widely adopted collaboration platform. The first half of May saw an even greater influx of newly identified issues. Engineering manager Hans Andersen implored his team to prioritize existing vulnerabilities, stating, "Please, please, please if your org has any April bugs, drive those down." He underscored the limited window of opportunity, emphasizing the need to "find as many things and do as much good as we can with this access" before May 31st, which he described as "the day when the rest of the world will have caught up."
This assertion sparked a sobering exchange among the engineers. One participant articulated the precariousness of their situation: "So basically you’re saying if it’s released on June 1, then on June 2 the adversaries will have our bugs?" The affirmative responses from multiple individuals highlighted the imminent threat landscape.
This situation aligns with predictions made by national security experts following Anthropic’s public unveiling of Project Glasswing in April. The initial hope was that the U.S. would possess a crucial window to fortify its software defenses before adversarial nations could deploy comparable AI models. However, evidence suggests this window may be rapidly closing, or has already shut. In late June, the Five Eyes intelligence alliance—comprising the U.S., Australia, Canada, New Zealand, and the U.K.—issued an unusual joint statement warning that this crucial period would diminish within months. The internal Microsoft meeting and related documents suggest that the "day of cyber reckoning" may have already arrived.
Triage in the Age of AI: A Shifting Risk Calculus
Microsoft’s current strategy, as revealed by internal records, involves prioritizing the patching of vulnerabilities classified as "critical" or "important," aligning with the company’s public patch updates. While plans exist to address "moderate"-severity flaws identified by Mythos, "low"-severity bugs were not mentioned in the reviewed documents. This approach mirrors the industry-standard triage system, akin to an emergency room prioritizing the most critically ill patients.
However, this established strategy faces a significant challenge in the current AI-driven bug-finding era. The sheer volume of vulnerabilities unearthed by advanced tools like Mythos introduces a new layer of risk. Mythos’s capability to "chain together" multiple bugs, where lower-severity flaws can be combined to create a significant exploit, means that seemingly minor unpatched vulnerabilities could become pathways for devastating attacks.
Vinh Nguyen, a senior technical advisor at Anthropic and a former chief AI officer and chief data scientist at the National Security Agency, voiced this concern. "The problem now is that you can chain four low-level flaws, and that can equal a high severity," he stated. "If you’re Microsoft, the current triage strategy may be underpricing risks."
In response to ProPublica’s inquiries, Microsoft defended its triage approach, asserting that decisions are based on multiple factors, including exploitability and customer impact. While the internal presentation did not explicitly mention bug chaining, a company spokesperson acknowledged that the technique "has long been considered as part of vulnerability assessment and risk analysis." The spokesperson downplayed the significance of the May 31st deadline, stating that "accelerated targeting and exploitation of new vulnerabilities is not a new phenomenon." Nevertheless, they acknowledged that the internal discussion reflected the company’s "sense of urgency to help our customers at this time" and reaffirmed that "security is Microsoft’s most important priority and teams across the company are prioritizing using AI to discover and remediate vulnerabilities as quickly as possible." Microsoft declined to provide figures on the number of bugs patched since the presentation. Anthropic also declined to comment.
The SharePoint Strain and a Growing Backlog
The internal Microsoft presentation projected that the team responsible for SharePoint would face a sustained workload, requiring "months" to address the backlog. This includes prioritizing critical bugs, followed by important ones in August. Microsoft defines critical vulnerabilities as those that can cause system crashes or facilitate malware propagation across networks. Important vulnerabilities can compromise the confidentiality, integrity, or availability of user data and processing resources. Following these high-priority fixes, the SharePoint team was slated to begin addressing approximately 300 "moderate" bugs.
While not a comprehensive overview of Microsoft’s entire product portfolio, the internal documents offer a glimpse into the magnitude of the challenge. Since its integration of Mythos earlier this year, Microsoft has collectively identified hundreds of critical and important bugs across popular products like Microsoft 365, the Teams conferencing platform, and the Copilot AI tool. As of mid-May, a significant portion of these remained unpatched.
"They’re not profound and exotic, but they’re real," remarked engineering manager Hans Andersen during the meeting, referring to the identified flaws. "And a lot of them are exploitable."
Escalating Patch Releases and the "Bug Apocalypse"
While it remains unconfirmed whether any specific bug identified by Mythos has been exploited by hackers, there is evidence of adversaries leveraging AI to automate attacks and identify vulnerabilities. This trend has manifested in Microsoft’s public "Patch Tuesday" releases. In June, the company issued patches for over 200 bugs, a number that industry experts at the time considered an all-time high. However, this record was shattered on July 14th when Microsoft released patches for more than 600 vulnerabilities. According to Dustin Childs, leader of the Zero Day Initiative bug bounty program at TrendAI, only seven of these were categorized as low- or moderate-severity, with one actively being exploited by hackers. The overwhelming majority were classified as important or critical.
"Well folks. Here we are. The bug apocalypse has fully descended upon us," Childs wrote in a blog post on July 14th, encapsulating the industry’s growing alarm.
Microsoft acknowledged to ProPublica that the overall volume of bugs "will not be plateauing for a bit." However, a spokesperson stated the company has "invested heavily in both people as well as AI-powered triage solutions that scale quickly to handle the growing number of vulnerabilities."
Rethinking Security in the AI Era
Experts like Vinh Nguyen advocate for a fundamental shift in how companies approach vulnerability management in the AI era. He suggests that instead of sidelining lower-risk flaws, organizations must dedicate resources to developing and testing patches for the entire spectrum of vulnerabilities, recognizing the potential for chaining. "There’s no alternative," Nguyen stated. "The patients are coming in fast and furious."
Microsoft indicated that it is "always going to be reevaluating and considering whether things that were previously lows or moderates be upgraded or thought about differently." The company acknowledged that "With these AI systems, it makes us rethink some of these things. Across the industry, we’re all looking to see how drastic of a change it will be."
Microsoft’s vast user base, spanning governments and businesses globally, makes it an attractive target for cybercriminals. Furthermore, many of its products incorporate "legacy" code—developed decades ago with outdated technology—which contributes to significant "technical debt" and harbors unaddressed flaws.
The challenge extends beyond Microsoft to the broader software industry, including the open-source community that underpins much of the world’s digital infrastructure. "Nobody has really figured out how to deal with this, and everybody is casting around for what they need to do," observed J. Michael Daniel, former cybersecurity advisor to President Barack Obama and president of the Cyber Threat Alliance. "Our tech debt is coming due."
Ben Edwards, a data scientist specializing in vulnerability management, described the situation starkly: "It was like drinking from a garden hose on the jet setting before, and now it’s like drinking from a fire hose. They might have had the teams that could handle that garden hose. Whether they can handle the fire hose is something else."
Understaffing and Corporate Priorities
Adding to the strain, Microsoft’s internal team responsible for handling vulnerabilities, the Microsoft Security Response Center (MSRC), has historically been understaffed. Even before the influx of AI-identified bugs, the center managed hundreds to thousands of reports monthly, pushing its resources to the limit. Former employees have suggested that this reflects a corporate philosophy where security is viewed as a cost center, while new product development is a profit center, leading to a reluctance to assign top engineering talent to patching efforts.
Microsoft stated it does not discuss internal staffing decisions but has recently made investments to "focus our teams on keeping our customers secure" and "continuously evaluates the staffing, processes, and technologies required to support security response and vulnerability management."
The internal presentation indicated that Anthropic provided Mythos access to approximately 50 full-time Microsoft employees with the explicit goal to "harden critical services before publicly available models catch up." A subsequent slide, "What’s Next," predicted continued high case volume for the MSRC "as public tools catch up" to Mythos.
During the May meeting, a staffer expressed a degree of comfort, suggesting adversaries "don’t have the source code." However, colleagues quickly corrected him, noting that portions of Microsoft’s source code have indeed been compromised over the years. "It might not be this week’s source code," one person responded, "But they’ve got source code. It’s out there." Microsoft later downplayed this remark, stating that its security processes are designed with the expectation that determined adversaries may gain access to code.
The race to secure software in the face of increasingly sophisticated AI-driven threats has become a critical battleground, with companies like Microsoft facing the daunting task of adapting their defenses to a rapidly evolving technological landscape. The implications of failing to keep pace could be far-reaching, impacting the security of individuals, businesses, and governments worldwide.







