The 60-second briefing

Microsoft introduced Project Perception on July 27, an agentic security system built around three roles: red-team agents find attack paths, blue-team agents investigate risk, and green-team agents take corrective action. Its first public-preview scenario, arriving August 3, pairs the MDASH vulnerability team with Microsoft’s MAI-Cyber-1-Flash model.

The useful number is 96% on CyberGym—12 points above Mythos—while Microsoft says the configuration cuts cost by almost 50% versus the current MDASH setup. This is a shift from AI that writes security alerts to AI that continuously reasons across identities, endpoints, applications, clouds, data, and other AI systems.

AWS published a task-aware knowledge-compression pattern on July 27 for the RAG problem that appears when an answer depends on hundreds of documents. Its reference design compresses the same knowledge base into task-specific versions at 8×, 16×, 32×, and 64× tiers instead of retrieving only the top matching fragments.

Build the new Cyber Stack

Microsoft’s design is a practical blueprint for builders: separate discovery, judgment, and action. A red agent can look for a vulnerability, a blue agent can rank whether it represents real risk, and a green agent can apply a fix. The important boundary is that the agents are specialized, while humans remain in control of the system’s permissions and governance.

The multi-model choice matters too. Microsoft says no single model optimizes quality, reliability, latency, and cost for every security task, so Project Perception chooses among frontier and specialized cyber models. For a smaller team, the same rule means routing cheap models to classification and expensive models to decisions that change production systems.

Cognizant announced an EMEA AI Unit on July 28 built around advisory, engineering, and delivery teams, including multi-agent delivery squads. The signal is practical: the market is packaging agent deployment as an operating service, not just selling another model endpoint.

AWS makes the same argument about context. Its example chunks documents into 256-token segments with 50-token overlap, compresses each chunk through Amazon Bedrock, stores the results in ElastiCache Serverless, and routes each query to a fidelity tier. The pipeline is concrete enough to reproduce: ingestion, compression, query analysis, retrieval, then inference.

The money is in context

AWS’s cost table uses a 100,000-token knowledge base queried 1,000 times per day. Full-context prompting consumes 100 million input tokens daily; its 32× “high” tier drops that to 3.125 million, and the 64× “ultra” tier drops it to 1.563 million. The savings come from paying the compression cost once during ingestion instead of shipping the entire corpus into every query.

That trade-off is not universal. AWS says the pattern fits knowledge bases that change infrequently, while hourly-changing data may still favor ordinary RAG. The right measurement is not “how small can the context get?” but whether the compressed answer preserves the numbers, entities, and relationships your task needs.

The labor signal is just as specific. OpenAI’s July 27 analysis of more than 800,000 messages from U.S. ChatGPT users found that 16.8% of work-related messages and 43.5% of occupation-specific messages involved tasks associated with another occupation. In workspaces with 2–5 seats, the outside-occupation share was 18.9%, compared with 16.3% in workspaces with more than 100 seats.

Try this workflow

Take one recurring process—security review, customer research, or a weekly report—and map it into three agents: discover, verify, act. Give the discoverer the broadest search permissions, require the verifier to preserve every date and number, and let the actor prepare a change without applying it. Add a human checkpoint before the final write.

❝

“Split this job into discover, verify, and act. Use original sources. Keep a list of every number, date, and named entity. Return the proposed action with evidence, then stop before changing files, sending messages, or publishing.”

Then measure two things: how many tokens the workflow sends after compression, and how many factual claims survive verification. That gives you the same two levers Microsoft and AWS are exposing this week—model selection and context economics—without pretending every task needs a frontier model or a fully autonomous agent.

— John