Malik Haidar has spent the better part of his career navigating the high-stakes trenches of global cybersecurity, serving as a primary architect of defense for some of the world’s largest multinational corporations. His approach has always been defined by a unique synthesis of deep technical intelligence and a pragmatic business perspective, ensuring that security is never just a checkbox but a foundational pillar of corporate strategy. As we stand at the precipice of a new era defined by frontier AI, Malik’s insights are particularly vital. We are currently witnessing a shift where the tools of the trade are no longer just assisting humans but are beginning to operate with a level of autonomy that was previously the stuff of science fiction.
The conversation today centers on the emergence of Claude Mythos and the strategic rollout of Project Glasswing—a massive, cross-industry initiative designed to harden the world’s digital infrastructure. We explore the profound implications of AI models that can unearth flaws that have remained hidden for nearly thirty years, the shifting economics of cybercrime which currently drains approximately $500 billion from the global economy every year, and the urgent necessity of a unified front between the private sector and government entities. Malik breaks down the transition from human-speed defense to AI-scale resilience, offering a look at how companies like Microsoft, Cisco, and AWS are already integrating these capabilities to stay ahead of increasingly sophisticated adversaries.
Frontier AI models like Claude Mythos are now autonomously identifying vulnerabilities that survived decades of human review and millions of automated tests. From your perspective in the field, what does this jump in capability do to our traditional understanding of a “hardened” system?
It completely shatters the illusion of safety we’ve built around legacy code. When you look at the fact that Mythos found a 27-year-old vulnerability in OpenBSD—an operating system that is practically the gold standard for security hardening—it tells you that our traditional methods have a massive blind spot. We’ve relied on the idea that if a piece of code survives for twenty years and passes five million automated tests, like the FFmpeg line did, it must be secure. But those millions of tests were essentially looking for what we already knew to look for. This model doesn’t just scan; it reasons. It autonomously identified and chained together multiple vulnerabilities in the Linux kernel to escalate from a standard user to full system control. That kind of sophisticated logic was previously the exclusive domain of the top 1% of human researchers. Now, it’s a capability that can be deployed at scale, which means we have to assume that every “hardened” system in our infrastructure likely contains critical flaws that are now discoverable in minutes rather than years.
Project Glasswing represents a massive coalition including Amazon, Google, and NVIDIA. Why is such a broad, multi-organizational effort required now, and what makes this different from previous industry collaborations?
The scale of the threat is simply too vast for any one company to patch the entire world’s attack surface. When you consider that open-source software makes up the vast majority of code in modern systems, a flaw in a single foundational library can compromise everyone from a small school to a major defense contractor. Project Glasswing is different because it’s not just a pledge; it’s an active deployment of resources. Anthropic is putting up $100 million in usage credits so that over 40 organizations can use Mythos to scan both their own proprietary code and the open-source infrastructure they rely on. We are seeing major players like AWS analyzing 400 trillion network flows daily, and they recognize that they need these frontier models to find threats before they emerge. This isn’t just about sharing information anymore; it’s about giving the defenders the same high-caliber “agentic” tools that adversaries are undoubtedly trying to develop. We are effectively trying to automate the defense of the entire internet simultaneously.
The financial cost of cybercrime is estimated at around $500 billion annually, with state-sponsored actors from nations like Russia and China posing constant threats. How does the introduction of AI-augmented exploitation change the economic and national security calculus for democratic states?
It creates a situation where the cost of an attack drops while the potential for destruction skyrockets. Historically, high-level exploits required months of work by highly paid, specialized experts. Now, the window between discovering a vulnerability and developing a sophisticated exploit has collapsed from months to mere minutes. For a state-sponsored actor, this is a force multiplier. It allows them to target critical infrastructure—power grids, banking systems, and medical records—with a frequency and precision that was previously impossible. This is why it’s a top national security priority for the U.S. and its allies to maintain a decisive lead. If we don’t use these models to find and fix the thousands of zero-day vulnerabilities in every major operating system and web browser right now, we are essentially leaving the door unlocked for adversaries who won’t have the same ethical safeguards we do. The economic damage of a single successful hit on a logistics network could easily dwarf our current defensive investments.
Open-source maintainers often operate on shoestring budgets despite securing the world’s most critical software. How significant is the $4 million in direct donations and the specialized access provided by this project for that specific community?
It is a game-changer because it addresses the “luxury” of security expertise. For a long time, only the Googles and Microsofts of the world could afford elite security teams to pore over every line of code. The average open-source maintainer was basically on their own. By donating $2.5 million to Alpha-Omega and OpenSSF and $1.5 million to the Apache Software Foundation, we are finally giving these “sidekicks” the tools they need to be proactive. These maintainers can now use Mythos-class models to find and fix bugs at a pace that matches the development of the software itself. It’s about democratizing high-end security. When a maintainer can use an AI to reproduce a vulnerability and generate a patch autonomously, we’re not just helping one person; we’re securing the millions of systems that run on that code. It’s the most efficient way to reduce the global attack surface because it fixes the problem at the source.
Looking at the performance data, Claude Mythos is outperforming previous models by significant margins, even achieving a 92.1% score on updated terminal benchmarks. What does this “agentic” reasoning allow the model to do that a standard coding assistant cannot?
The difference lies in the ability to handle multi-step, complex problem-solving without a human holding its hand. A standard assistant might help you write a function, but an agentic model like Mythos can look at a binary, perform black-box testing, identify a memory corruption issue, and then figure out how to bypass modern mitigations to achieve execution. On the Terminal-Bench 2.0 tests, we saw it navigating environments and using tools with a level of persistence that is frankly startling. It can use 4.9 times fewer tokens than earlier models like Opus 4.6 while achieving better results, which shows an incredible increase in reasoning efficiency. This means it can “think” through a security problem, test its own hypotheses, and iterate until it finds the flaw. In a defensive context, this allows us to automate penetration testing at a scale that was previously unthinkable. We can now have an “AI red team” running 24/7 against our foundational systems, constantly trying to break things so we can fix them first.
What is your forecast for the state of global digital infrastructure over the next few years as these frontier models become the primary tools for both defense and offense?
I believe we are entering a period of “hyper-patching.” In the near term, we are going to see a massive spike in reported vulnerabilities—thousands of them—as Project Glasswing partners unleash these models on their stacks. This will be a chaotic but necessary phase where we flush out decades of technical debt and security flaws. Within 90 days, we’ll start seeing the first public reports on the vulnerabilities fixed, and that will set a new baseline for what “secure” looks like. However, the real shift will be in how we build software from the ground up. We will move toward a “secure-by-design” lifecycle where AI models are integrated into every stage of development, catching bugs before they even reach a repository. The “minutes vs. months” reality means that our update processes and triage scaling have to be fully automated. If we succeed, we will eventually reach a state where the cost to an attacker is so high, and the defense is so fast, that the current $500 billion drain on our economy begins to actually shrink for the first time in the modern era. We have a narrow window to act, but for the first time, the defenders actually have the tools to win.

