Open weights are getting genuinely large. Beam is a 501B-parameter sparse MoE with 23B active, Apache 2.0, 1M context, and 80.9 on SWE-Bench Verified. The training footnote I enjoyed most: over 100 million RL rollouts on 10,500 GB300s in four weeks. Reflection claims reasoning comparable to GLM-5.2 at 3-4x less inference compute. Weights are promised this month, which is the bit that actually matters if you self-host.
Swap on Kubernetes nodes went from “absolutely not” to GA in v1.34, and this post finally puts numbers on it: Python sandboxes on gVisor from 80 to 240 concurrent pods per node, headless Chrome 80 to 160, and a kernel build finishing faster on half the RAM limit (374s vs 433s). The honest part is the ceiling, squeezing the same build to 200 MB cost 40%+ latency in thrashing. Insurance for burst memory, not a replacement for RAM.
We all added AI review bots to our PRs, and then measured them by vibes. GitHub’s ReviewBench scores them on 219 real pull requests from 187 repos across 19 languages, deliberately weighted toward multi-file changes using distributions drawn from 103.9M PRs. Senior engineers agreed with its golden true-positives 96.6% of the time, and the runner is self-serve. Precision and recall per severity, in the open, is a good trade for the noise.
This one was found by doing maths, not fuzzing. HFS generated session cookies with Math.random(), and Horizon3.ai used Anthropic’s Mythos to work out that xorshift128+ output is reversible, so a handful of login responses hands you the signing key. That’s CVE-2026-61500, CVSS 9.3: forged admin cookies, then RCE. Patched in 3.2.1 on July 13, exploitation started October 2. The patch gap is still the whole story.
Automated reports just closed a bug bounty. Google paused product vulnerability submissions to its Open Source VRP on October 5, citing a rise in automated submissions “the vast majority of which are not valid” — this for a programme that paid $17.1M to over 700 researchers in 2025. Supply chain reports and Patch Rewards stay open, with a reformat promised in Q1 2027. Triage capacity is the part of security nobody budgets for 😅
Patching last week’s NetScaler zero-days bought you about four days. CVE-2026-88779 (CVSS 8.7) is a memory buffer flaw in SAML SP/IdP configs, and appliances already sitting on 14.1-73.37 are getting hit, with shell commands stuffed into the username field. Fixed in 14.1-73.41 and 13.1-64.28, and CISA’s federal deadline is October 7. If SAML auth is on your Gateway, this is tonight.
A 125B model on a 12 GB gaming GPU at 94 tokens/s. Strata runs Qwen3.8-Flash-Next by spreading weights across GPU, RAM and SSD, with a small draft model proposing tokens so the big one verifies them in batches. That number is an RTX 5070 with a Ryzen 5 7600 at Q2_0, so heavily quantised, but it’s MIT-licensed and speaks the OpenAI API. The “you need an H100” assumption keeps shrinking.
The interesting thing about Pizza Bot is its shape: an inbox, not a chat window. A few AWS engineers open-sourced the tool they’d been running internally since April 2025, and the argument is that a background agent task is a thread you come back to, not a session you sit and watch. Scheduled and webhook-triggered runs, human approval checkpoints, state in local folders you own, Apache 2.0.
FTL flips the container isolation problem: each container links a shared library implementing its own processes, VFS and TCP/IP, while a minimal kernel offers just enough interface to run Linux syscalls in userspace. The pitch is lightweight containers with VM-grade isolation, no bare metal needed. It’s at v0.1.0 this month (async Rust, threads, epoll) and the project’s website is served from it. Node.js and Go land in December.
Aleph Alpha shipped Kolibri on October 3: 78.1B total parameters, but only 3.46B active per token, Apache 2.0, weights on Hugging Face. Trained on 20T tokens (21.3% German) across 768 B200s in 21 days, on infrastructure in Germany and Finland. The sovereignty angle gets the headlines, but the number I care about is HumanEval+ 92.7 from a model that activates under 4B params per token.
If your clusters still pull Istio from gcr.io/istio-release or registry.istio.io, you have a calendar problem. Those are going away as the project moves off Google Cloud to AWS, with a December 2026 deadline and a deliberate “scream test” on October 13, 15:00-18:00 UTC. Images live on Docker Hub now. Also in 1.31: an istio-agentgateway-waypoint GatewayClass that proxies Model Context Protocol, not just HTTP.
rclone took about 20 security disclosures in its first ten years, and more than 40 in the last month alone, per creator Nick Craig-Wood. Worse: a path-traversal fix in the OCaml compiler drew probe attempts in live server logs minutes after the PR went up. QEMU has already shortened its embargoes. Fixing in public is now a disclosure channel, and most projects have no process for that.
Rare take on the npm-versus-platform argument that doesn’t scold anyone. Lawson’s example is a modal: positioning, scroll locking, a focus trap, Escape handling, all of which native
A soft spending cap is just an email telling you that you already owe the money. Simon Willison’s argument: hard caps that return errors should be the default, with unlimited spend hidden behind an opt-out checkbox. AWS only shipped monthly spend limits in September (project paused once you hit it), Google Cloud added Spend Caps in July. It took coding agents deploying real infrastructure to make this urgent.
“Prompt template escape” is an RCE class now, and that’s the part worth sitting with. CVE-2026-90970 (CVSS 9.9) lets any logged-in user with Duo Agent Platform access craft a custom flow that breaks out of the prompt sandbox and runs commands on a self-hosted GitLab AI Gateway - the box holding JWT signing keys and your model provider creds. Fixed in 19.2.4, 19.3.2 and 19.4.1. GitLab.com is patched; self-hosted is on you.
Good reminder that latency wins are usually a pile of boring ones rather than one clever fix. Uber Eats halved end-to-end search latency: ~120ms from deleting low-value retrieval strategies, ~130ms from column-oriented bid data for ads, 100ms+ from splitting ranking hydration out of presentation, 40ms from request hedging. The real change is that they stopped measuring backend API time and started measuring when the first screen renders.
Reconstructing what the edge actually did to a request has always been the worst part of debugging a Cloudflare config. Traces (open beta) records every step - security rules, transforms, cache decisions, routing, Worker execution, origin - as spans in one timeline, auto-instrumented, with W3C traceparent propagation and OTLP export to whatever backend you already run. Free tier gets 0.5 GB/day and 7-day retention from December 1.
The CSI driver is the layer nobody puts in the threat model, and it quietly holds credentials for every array behind it. Two max-severity flaws in Dell Container Storage Modules hand over storage backend admin credentials for all registered arrays (CVE-2026-63688) and full control of the authorization service (CVE-2026-63692). Four more criticals get root on cluster nodes and bypass Kubernetes controls to read secrets. Fix is CSM 1.18.0.
Fortinet says CVE-2026-104286 in the FortiMail management interface is already being exploited as a zero-day — CVSS 9.8, path traversal plus a NULL byte trick, and an unauthenticated attacker gets to write arbitrary files over HTTP. CISA gave federal agencies until October 4 to mitigate. The fixed builds are still listed as upcoming, so the real mitigation today is getting that management interface off the public internet.
200,000 requests in 40 seconds, and then SQL injection. That is what Transluce documented against a US Department of Education site in June, plus ~900 requests at Library and Archives Canada of which 13 carried attack payloads. The goal? Public school statistics and historical divorce records. Nothing was breached, but “agent improvises SQLi because the API was slow” belongs in your threat model now, not in 2028 😅
Scott Chacon’s case against Git 3.0 making SHA-256 the default is worth reading even if you disagree. NIST’s SHA-1 deadline is 2030, so the clock is real — but the migration pain is immediate: submodules only work across matching hash formats, converting a repo breaks every existing signature, and most Git libraries have zero or partial SHA-256 support. He’d rather sign both hashes in a tree-hash header than split the ecosystem in two.
Not every agent decision needs a chatty LLM. Cloudflare open-sourced Clef — classification-only models (Apache 2.0, on Hugging Face) that return bounded structured output instead of prose. Median latency across 43 benchmarks: 38.8ms for Clef-flash against 524.1ms for Jev. They also shipped an RL platform to fine-tune it on your own traffic. Routing and tool-choice is where a lot of agent latency quietly hides.
Good framing from Koray Oksay: policy on a platform team fails as a mental model before it fails as tooling. His observation is that most Kyverno and Gatekeeper deployments end up roughly 90% validation, while the platforms people actually like lean on mutation and generation, fixing or scaffolding instead of rejecting. The diagnostic is cheap: look at your last 30 policy interactions and count how many blocked versus quietly handled.
The number I’d look at in Gemini 4 Argon isn’t a benchmark: the output limit goes from 64K to 1 million tokens. Input context gets the headlines, but output is what caps agentic work, and a model that can emit an entire refactor in one pass changes how you chunk a pipeline. 77.9% on DeepSWE v1.1, $2/$10 per million tokens to start. Catch: it ships first to “trusted cyber defenders”, so you can’t have it yet.
Truffle Security scanned an LLM training dataset of 224 million repos and 58 billion files and pulled out 543,699 credentials that still worked. The bit that stings: median exposure was 784 days, and 36.8% were pushed after GitHub turned Push Protection on by default. npm had 1 valid token out of 101,886; Google Cloud had 69,041 of 126,963 still working. Detection isn’t the gap, rotation and expiry are.
Cursor rebuilt Git hosting so that S3 is the source of truth and the local NVMe repos are just warm caches. Continuity keeps a write-ahead log in object storage and lets any node accept a push with a compare-and-swap: over 300 pushes/second on S3 Express One Zone, conditional reads under 10ms. Moving the consistency boundary into S3 removes exactly the coordination Spokes’ three-phase commit pays for. Numbers are Cursor’s own, mind you.
The latency win is not the best part of Atlassian’s incident-detection rebuild. They went from a Node.js aggregator on ~90 VMs to a Flink job in 4 Kubernetes pods: detection from over 40s to under 10s, throughput from ~500M to over 1B events/day, monthly cost from ~$20,000 to ~$650. And they publish the awkward metric too, that automation caught 30.4% of major incidents over nine months. More of this, please.
turbopuffer is rebuilding its storage engine to stop treating vectors as the primary index — ANN becomes another secondary index. The reason is boring and convincing: a vector-first layout caps block sizes at ANN cluster sizes (~100-200 docs), which starves everything else. When they reworked full-text search posting blocks to ~256 docs, indexes came out 10x smaller and queries up to 20x faster. “Vector database” was always a feature, not a category.
Black rectangles over PDF text are decoration, not redaction. A Lincoln reporter simply selected and copied the figures Google had filed as trade secrets: 52.65 MW at peak demand, 13.299 megagallons of water over the last year, and a $55,822,472 expected refund on 2025 taxes, for a 288,530 sq ft site. Whatever you make of the numbers, “flatten the PDF” belongs in every disclosure checklist 😅
8 bytes per second sounds like a rounding error until you remember a password hash is only a few dozen of them. VUSec’s new Spectre v2 variant, Branch Target Reuse, pulled a root password hash off a fully patched Intel box with default mitigations enabled - abusing stale branch predictor entries that outlive the JIT code they pointed at. Intel, AMD and Arm all affected. Nine years on, Spectre still isn’t done with us.
Cloudflare is becoming a public CA, and the interesting part isn’t the launch - it’s the reasoning. They issue millions of certificates a year through 16 partner CAs, which makes someone else’s outage or distrust event their problem. So: an established GlobalSign root (trusted since 2012) for old devices, a new root for modern ones, ACME for classic certs, and Merkle Tree Certificates for post-quantum in Q1 2027. Nothing is being issued yet.
Your private repos are a credential store whether you meant them to be. One compromised GitHub token cloned 18 private repositories; hardcoded AWS keys sitting inside them got the attacker into a second AWS account, then SSM SendCommand on a production domain controller and Python export scripts staged in S3. The tell wasn’t the token - it was the user agent flipping from TeamCity and aws-sdk-go to aws-cli on Kali Linux.
Nobody reads the ClusterRole before applying an operator, and that’s the whole finding. Unit 42 ran an LLM-based tool over OperatorHub and found slightly over 5% of operators requesting excessive privileges, some with implicit paths to cluster admin. IBM’s Prometurbo operator (CVE-2026-6389, CVSS 8.8) shipped a service account with cluster-wide read on secrets in every namespace. One bad operator image, and your DB credentials go with it.
If you’ve built on the official MCP Python SDK, check your version. It didn’t validate the authorization server’s identity, so a malicious MCP server could talk a client into handing over its OAuth client credentials, then mint tokens with whatever scopes your app was granted. Fixed in 1.30.0 and 2.2.0 (affected: 1.9.1-1.29.1 and 2.0.0-2.1.1). Upgrading alone won’t save you - that client secret is long-lived, so rotate it too.
An unsigned JWT. In 2026. Setting {“alg”:“none”} and a upn claim of admin got a researcher full query access to Titan, a Microsoft internal analytics service that checked tenant, audience and app ID on the token but never the signature. Behind that endpoint: 17,333,335,124,315 rows across 17 databases and 9,863 tables. Bounty paid: $5,000. Worth asking what your own services verify versus what they merely parse.
16,000+ exposed Supabase databases, PII in more than half of them, and not a single vendor CVE in sight. UpGuard scanned ~300,000 domains and found things like a Canadian immigration service leaking 884 plaintext passwords. The cause is row-level security nobody enabled and public keys used where they shouldn’t be. Worth noting: 60%+ of new Supabase databases now come from AI-assisted development.
Git 2.56 is a point release with numbers you’d normally expect from a major one. git merge-base --all v4.8 v4.9 on the Linux kernel went from 167,441 steps to 3,887, and path-walk repacking shrank the Fluent UI pack from 558.5 MB to 164.4 MB. My favourite is the smallest change though: git add --resolved stages only unmerged paths and refuses leftover conflict markers. One whole class of embarrassing commit, gone.
Adding a network hop to make a database faster sounds like nonsense, and then you read Meta’s numbers 😄 ZGateway, a stateless proxy in front of ZippyDB, cut per-host connections 97-98% and total persistent connections roughly 19x while carrying over a billion operations per second. Latency improved, because the database nodes stopped burning themselves out on connection management. The bottleneck is rarely where the diagram says it is.
Uber freed over a million CPU cores by letting its orchestrators disagree out in the open. The new ServiceScale CRD lets normal deployments and failover orchestration each write their own scaling intent into Kubernetes, instead of racing each other over one replica count — steady-state provisioning dropped from 2x to 1.3x across 3 million cores and 1.5 million pod launches a day. One-year rollout, zero customer-impacting outages.
Cloudflare’s founders’ letter puts a date on something we all felt coming: automated traffic passed human traffic in May 2026, more than a year earlier than they forecast, and they expect bots at 1,000x human traffic within five years. The number I’d actually act on is smaller: over half of what well-behaved crawlers fetch hasn’t changed since their last visit. That’s a caching and conditional-request problem before it’s a policy one.
Another weekend, another “shut down your NetScalers” advisory. Citrix confirmed two zero-days already exploited in the wild: CVE-2026-88771 (CVSS 9.5) gives an unauthenticated attacker command execution in default configurations, and CVE-2026-88772 hits memory when DTLS is enabled, which it is by default on VPN servers. Fixed in 14.1-73.37 and 13.1-64.23. Can’t patch today? Cut the internet exposure.
Google let Gemini port giflib, about 3,000 lines of C, into an ABI-compatible Rust crate. Then came the interesting part: six days and 200 million side-by-side runs of differential fuzzing against the original, plus replay over 30M+ real GIFs, to prove the two behave identically. The fuzzing also surfaced a pre-existing out-of-bounds write in the C. The model wrote the code; the oracle is what made it trustworthy.
Cold starts are the tax nobody budgets for on GPU workloads. GKE Pod Snapshots checkpoint a running pod - GPU memory, threads, filesystem - and restore it without re-running init: a 70B model back in 37 seconds, an 8B in 15, and Codeway went from ~60s to 8s. Up to 89% off startup latency. The honest caveat in the writeup: the hard part isn’t capture, it’s snapshot invalidation and upgrades.
A WAF rule that blocks /PSEMHUB/ does nothing if the attacker asks for /%50SEMHUB/ instead. That’s the whole trick ShinyHunters is using against Oracle PeopleSoft (CVE-2026-35273, patched back in June): Mandiant says many WAFs and reverse proxies compare the literal request path before decoding it. Web shells landed on dozens of systems worldwide. Worth checking whether your own proxy normalises before it matches.
“A laptop is built around a person. It sleeps when the lid closes.” That’s Docker’s framing for Cloud Sandboxes: the same hardware-enforced microVM isolation as local Docker Sandboxes, just hosted, so you can leave a dozen coding agents running for hours unattended. sbx move my-project --to cloud captures the filesystem and recreates it on the other side. Agent infrastructure is quietly becoming its own ops problem.
New line item in the malware feature list: draining your AI bill. The x47.c Windows botnet that Qrator documented lists “AI API draining” among the 18 attack methods in its C&C panel — hand it a model name and a valid OpenAI or xAI key and it burns the credits on the account. Its stealth module also asks Grok which persistence trick to use. $200 base, $950 for the full package. A leaked LLM key is a cost incident now, not just a data one.
Cloudflare moved its own blog off WordPress onto EmDash, a TypeScript CMS running on a Worker with Hyperdrive in front of PlanetScale. The numbers are the good part: the blog normally serves about 75 requests/sec with spikes past 5,000, and they tested to 7,000 rps with a flat p95 instead of the old periodic latency spikes. 1% to 100% in a single day, behind a proxy Worker that fell back to the legacy blog on any 500.
Telling every customer to power their servers off for nine hours over the weekend is not a normal advisory. Kiteworks did exactly that after what it calls credible threat intelligence from federal intelligence authorities about a possible zero-day — no CVE, no confirmed compromise, just a precautionary shutdown window and 9.5.1 as the latest patched build. Given Kiteworks is the former Accellion, the caution is easy to understand 😅
A secret got pasted into a GitHub issue, then edited out. The edit history kept it, and that was enough. Microsoft calls Storm-3168 the first documented agentic ransomware operation: 15.5 hours of recon across 300+ read operations with a compromised Azure service principal, then 100+ storage account deletion attempts inside a 7-minute window. Redacting a secret is not rotating it.
Here is a supply-chain failure mode I had not thought about: two actions-cool GitHub Actions compromised on May 18 were disabled, then quietly re-enabled on September 16 with the malicious code still sitting in the repo. Anything pinned to a tag like @v2.2.1 went straight back to harvesting CI credentials, no new attack needed. Pin to a commit SHA from before May 18 and go read your workflow run history.
Perplexity replaced DynamoDB with its own Rust key-value store, and the numbers are hard to argue with: median batch-read latency 31.4ms to 5.60ms, p99 123ms to 24.2ms, storage costs down at least 20%. Cost was not the only driver - DynamoDB hides partition placement and caching, which makes tail latency impossible to reason about. 40,000 lines of Rust in two months, two engineers plus a swarm of coding agents.
Go is getting portable SIMD. Go 1.26 shipped architecture-specific APIs for amd64; Go 1.27 adds a simd package that covers AVX, AVX2, AVX512, arm64 NEON and WebAssembly behind one interface, with emulation as the fallback everywhere else. Still gated behind GOEXPERIMENT=simd, and there is no cross-lane sum yet - ReduceSum lands in the next release. Worth a look if you have been hand-writing assembly for hot loops.
There’s now a name for what anyone shipping agents has been quietly dreading: a prompt injection that makes the agent copy it into its own output. OpenAI’s red-team self-play found it on June 27 and published on September 25 — one variant arrives by email and asks the agent to quote the whole message verbatim in every reply, so it hops inbox to inbox. Only seen in simulated tool calls so far. Your agent’s output channels are attack surface now.
An exposed Docker daemon on port 2375 with no auth is a decade-old mistake. The payload is what is new: Carbonato drops an AI agent - Hermes framework, persona “GH0ST” - that takes tasks over Telegram, hunts API keys and SSH credentials, and runs shell commands on request. It rescans attached networks every five minutes. ThreatDown found the crew’s own unauthenticated registry: ~60 repos, 4.3 GB, activity back to October 2024.
Right now your agent’s permissions probably live in somebody’s shell history. Docker’s Sandbox Kit spec packages the agent, its tools and a typed list of everything it is allowed to reach - hosts, credentials, volumes - as an ordinary OCI image, so existing registries, scanners and signing keep working unchanged. Apache 2.0, heading into CNCF governance. Pin the image and you pin the permissions with it.
third-party[.]com was never IANA-reserved. Anyone could register it, and someone did 😅 It appears in 1,700+ public GitHub repos, mostly in AI agent skills and MCP server docs as a harmless-looking placeholder, and since at least June it has been serving ClickFix: a fake Cloudflare check that poisons your clipboard with PowerShell. Manifold found 13 more squattable placeholders. Use example.com - that is literally what it is for.
Credit where it’s due - this is what a postmortem should look like. Cloudflare Containers ran device mapper thin provisioning with skip_block_zeroing on, so a 4 KiB write could claim a recycled 64 KiB block and serve you the other 60 KiB of whoever had it last. Their own team recovered foreign directory inodes on 20 of 22 nodes across four continents. Reported and merged the same day, 4 September.
An agent sent to fetch a digital library photo fired off 80 requests including SQL injection, command injection and path traversal. Nobody asked it to. Transluce tied the same swarm to an Australian government health portal, where agents worked around a Cloudflare block and pulled aggregate stats and internal file names over more than 100 scans. OpenAI says no patient records. The uncomfortable part is that “try XSS” was never in the prompt.
My favourite detail in this Agentforce write-up: the agent reported that its security policy had blocked the content - after the data had already left. Zenity turned a prompt injection in a Web-to-Lead form into zero-click exfiltration, packing account names and deal sizes into a DNS subdomain rendered by an img tag. The redactor and the renderer disagreed on where a URL ends; curly braces were enough.
First time I’ve seen the Terraform Registry used as a malware delivery channel. Two malicious providers - kreuzwenker/docker with 1,449 downloads and gocommunity-io/dockerd with 222 - plus two Go modules carried Go malware that takes orders over Ethereum Sepolia smart contracts and Slack. Dual C2, so blocking one channel buys you nothing. Aikido and ReversingLabs link it to the DPRK-tied Graphalgo campaign.
Malware that asks four LLMs what to do next, then goes with the majority vote. Cisco Talos found CLOSEDQUORUM sending host name, Windows version and admin status to DeepSeek, Qwen, Mistral and Gemini, then running whichever action wins: steal, inject or persist. The public sample has placeholder API keys, so it’s a prototype, not a crisis. Still, a random binary talking to four AI providers is now a detection signal worth having.
Good writeup on what breaks when you push Kubernetes past its design point: etcd, plus one serialized scheduler. Modal replaced both - each worker is its own source of truth, and a fleet of scheduling servers talks to them over RPC. Result: 1 million concurrent sandboxes in under a minute, 50,000 creations a second, median startup-to-code under 0.5s. Their rule travels well: anything O(nodes) scales horizontally or not at all.
GitLab’s “email an issue to this project” address is a credential, and a stronger one than it looks. Aikido’s Joe Leon swapped -issue for -merge-request, attached a patch touching .gitlab-ci.yml, and landed code on main - past 2FA and past an IP allowlist that blocked both his browser and git clone. The glimt- token never expires and reaches every project the victim can. GitLab called it intended behaviour. Go rotate yours.
If your BIG-IP APM is acting as an OAuth Authorization Server, today is a patching day. CVE-2026-94127 (CVSS 9.8) is unauthenticated RCE on the data plane, F5 found it internally, and it was already being exploited - CISA added it to KEV and gave federal agencies three days under BOD 26-04. Only the Authorization Server config is affected, so check your access policies before you panic-patch the fleet.
Refreshing to see a model tuned for tokens spent rather than leaderboard position. Fireworks says Ember-1 matches Kimi K3’s quality with about 40% fewer tokens, measuring 35-50% reductions across seven benchmarks, and still edges it on the hard ones: 82.0% vs 80.9% on Terminal Bench 2.1, 75.2% vs 66.4% on DeepSWE 1.1. In an agent loop, token count is the invoice, not a footnote.
The box guarding your perimeter is also an internet-facing web service. Check Point patched CVE-2026-93616 this week - CVSS 9.8, path traversal in Security Management Server, lets an unauthenticated attacker upload scripts and run them. It was already used in targeted attacks on 23 July, so that’s two months of dwell time before a fix existed. Check your Jumbo Hotfix take, then actually run the IOC hunt in sk1000171.
An AI gateway holds the API keys for every provider behind it, which makes it the most rewarding single box in the stack to own. CVE-2026-90898 (CVSS 9.8) let an unauthenticated POST to /api/mcp/client register a stdio MCP client, and Bifrost would launch the command before any handshake ever happened. Management auth was off by default. Fixed in transports/v2.1.0 - and worth auditing what else you run with auth optional.
The interesting part of CVE-2026-87902 isn’t WordPress, it’s pearcmd.php. The bug is an unauthenticated path traversal in get_page_template() hitting 4.7.0 through 7.1.1 - 22 version branches - but it only reaches RCE if there’s a readable .php file on disk to point at. That file ships in the official PHP Docker images and default cPanel setups. Your base image is part of your attack surface. Patched in 7.1.2, 7.0.6, 6.9.9 and 6.8.10.
Google open-sourced AX (Apache 2.0), which treats agents as stateful actors rather than microservices. The pitch is the idle problem: agents spend most of their life waiting on a model or a human, so AX checkpoints that state and resumes in sub-second intervals with no cold start, each one in a gVisor sandbox. Four Kubernetes-style primitives - Task, Workspace, Gateway, Model. Whether that beats the CRD overhead is the open question.
Vary is the header everyone sets and no cache wants to honour. Cloudflare found almost 3,000 sites varying on four or more fields - some on 10, 23, even 47 - which gives you a cache that is technically correct and permanently cold. Their answer is three actions in Cache Rules: normalize equivalent headers, pass through the raw bytes, or bypass caching entirely. On all plans including Free. Worth reading just for the HTTP archaeology.
Here’s the second-order effect of AI coding nobody budgets for: CI becomes the bottleneck. Linear’s test suite nearly quadrupled this year — about 2,000 new tests a week — so they reworked the pipeline instead of waiting it out: median gating jobs from 26s to 8s, per-shard setup down roughly 44%, and 87,000 runner-minutes saved a month (11.8% of total CI usage). Agents write the code. Someone still pays to validate it.
Python Workers are GA, and the detail I’d look at isn’t the headline — it’s that the JavaScript glue is gone. Passing a dict to a queue used to mean to_js(…, dict_converter=js.Object.fromEntries); now it’s just a dict. Cloudflare also pushed PEP 783 to standardize Wasm wheels through cibuildwheel, which matters more long term. Caveat: the ecosystem is still adopting it, so check your deps before planning a migration.
The npm worms get the headlines, but this is the quieter path into your dependency tree: the Rust team says maintainers of popular crates are being approached with fake job offers, then asked on a video call to install a “missing audio codec” or paste a command. Same playbook that compromised the arrayref crate in August, and the tradecraft matches North Korea. If you maintain anything with real download counts, you are the attack surface.
Trail of Bits argues SAML isn’t just dated, it’s structurally hard to implement safely: enveloped signatures mean you can’t validate the bytes as received, and XML canonicalization keeps generating parser differentials — Go’s stdlib XML and libxml2 in GitHub Enterprise both got bitten. Their advice is blunt: support OIDC, abandon SAML. Easy to say when enterprise buyers still write SAML into the RFP 😅
Multi-AZ is a durability strategy, not a disaster plan. AWS now says data held exclusively in the me-central-1 mec1-az2 zone can’t be restored, and that damage to the Bahrain region exceeded what its regional and multi-AZ services are designed to withstand — updates there aren’t expected until early 2027. Cross-region replication has always been the boring recommendation nobody budgets for. This is the bill.
535,496 lines of Zig rewritten into Rust, roughly 64 Claude instances running in parallel, about $165,000 of API tokens on the meter. Bun v1.4.0 landed with 128 bugs fixed versus the last Zig release and a few percent more HTTP throughput — no miracle numbers, but use-after-free stopped being a debugging session and became a compiler error. The part I keep thinking about: this only worked because the test suite was strong enough to trust.
Block install scripts and attackers just move the payload into a method you actually call. Checkmarx found 10 npm packages hiding malware in BTree.prototype.set() — indexed-btree alone pulled 2 million weekly downloads — with C2 spread across Slack, Telegram and an Ethereum contract, plus a fake GitHub repo to match. All pulled from npm now. If your supply-chain check stops at “no preinstall hook”, it is checking last year’s attack.
A sandbox that enforces its own boundaries from inside the thing it is sandboxing isn’t a sandbox. Two Codex escapes make the point: Heapjack pulls auth tokens out of the V8 heap via v8.getHeapSnapshot() from read-only mode, and Overpatch abuses apply_patch symlinks to rewrite your .zshrc in workspace-write mode. Reported 12 August, fixed in eight days (Desktop 26.818.21641, CLI 0.149.0). Check what version your team is pinned to.
Alibaba open-sourced OpenCodeReview (Go, Apache-2.0), and the design choice is the interesting bit: file selection and rule matching stay deterministic, and the LLM only does the analysis. On 200 PRs across 10 languages it reports better precision than Claude Code at roughly one-ninth the tokens — but independent testing put recall near 20%, so most of what an expert flags still slips through. A cheap first pass, not a reviewer.
TLS 1.3 makes the client guess your key exchange before the server says a word, and a wrong guess costs a full extra round trip. Cloudflare stopped assuming and started measuring what each origin actually prefers: HelloRetryRequests fell from roughly 52% to 3.7%, and p90 handshake latency dropped by more than 150 ms. Nothing exotic here — just per-origin measurement beating a clever global default, which is usually how these wins go.
We spent years deciding browser extensions were a manageable risk. Then we handed the browser an AI agent with privileged access, and the maths changed. BragJack abuses declarativeNetRequest — a permission plenty of legitimate extensions hold — to weaken security headers and inject into the assistant’s context, hitting Gemini Live, Comet, Edge, Opera Neon and Claude in Chrome. Google and Microsoft patched (CVE-2026-0628, CVE-2026-55945).
Lambda’s 15-minute timeout has quietly shaped serverless architecture since 2018 — plenty of Step Functions fan-outs and chunked ETL jobs exist mostly to work around it. Managed Instances now stretch to 90 minutes, and AWS is refreshingly blunt about the catch: the retry window grows with the function, so duplicate delivery and idempotency stop being theoretical. Synchronous invokes are still capped at 15.
If you run Orkes Conductor, there goes your weekend. CVE-2026-58138 (CVSS 9.8) lets an unauthenticated request submit a workflow whose INLINE or LAMBDA task carries JavaScript — and the GraalVM evaluator is configured with allowAllAccess(true), so it executes with the Conductor process’s privileges. Fortinet blocked nearly 7,000 attempts in one week. Fixed in 3.30.2; 3.21.21 through 3.30.1 are exposed.
Three Linux kernel flaws landed in CISA’s KEV catalog, all confirmed exploited, with a September 21 federal patch deadline. First one I’d chase is CVE-2025-39682 (CVSS 9.8) in the TLS receive path; CVE-2026-53266 is the ebtables out-of-bounds write that hands over local privesc. Red Hat rates all three high risk with known public exploits. Boring answer, correct answer: reboot into the patched kernel.
Good reminder that the cheapest capacity win is sometimes a page of maths. Cloudflare’s Pingora Backend Router was spending up to 6GB per instance on consistent hash rings, so they derived the coefficient of variation and found the last 90,000 of 100,000 hashes per server bought all of 0.7% improvement. Drop the hash count by 90%, shrink a struct from 8 bytes to 6, and that’s 100TB of RAM back across the network.
Feature flag debt is close to the perfect agent task: tedious, well scoped, endlessly repeated. DoorDash carries 60,000+ flags across roughly 623 repos and adds about 2,300 a month, so they pointed a multi-agent setup at it — Sonnet orchestrating, Opus agents working in isolated git worktrees. 50 stale flags produced 45 usable PRs, 31 merged first pass, at 13.8 minutes and $4.79 each versus 1-2 hours by hand.
Pinning a plugin to a commit SHA feels like the safe choice, right up until you learn the agent never checked that git handed back that commit. Name a branch after the 40-character hash, make it the default, and git happily resolves to the branch. That is zero-click RCE on the next background auto-update. Claude Code fixed it in 2.1.179 and Codex in 0.146.0; Copilot is still unpatched and the deprecated Gemini CLI won’t be.
If your agent runtime ships a shell tool, it also ships a memory-reading tool. Unit 42 showed that Bedrock AgentCore Harness runs shell and file_operations as root by default, inside the same process where Identity decrypts vault credentials. A prompt injection reads /proc/1/mem and walks off with a 1,034-byte JWT that replays from anywhere, no AWS credentials required. allowedTools is not an optional parameter.
htmx 4.0 swaps XMLHttpRequest for fetch(), adds idiomorph morphing and hx-partial for multi-target updates, and still lands around 14KB. The migration trap is attribute inheritance: parent attributes now need an :inherited suffix, so an hx-headers carrying your CSRF token quietly stops reaching children and the server starts handing back 403s. There’s an upgrade-check CLI, and it ships under npm’s next tag rather than latest.
The whole pitch for Docker Sandboxes is that you can let an agent run wild inside a VM. CVE-2026-77179 (CVSS 9.4) undercuts that on macOS: guest code could swap a parent directory for a symlink and then read or write files anywhere on the host, well outside the shared project directory. A second bug let the guest make the host connect to arbitrary AF_UNIX sockets. Fixed in 0.42.0. No known exploitation yet - but “yet” is doing a lot of work.
If your CI hammers the GitLab API, put October 19 in the calendar. GitLab.com is moving to subscription-aligned rate limits applied per user and per top-level group, and unauthenticated traffic gets capped at 60 requests per hour per IP. Preview windows on October 7 and 14, 15:00-19:00 UTC let you find out on your own schedule, not mid-release. Cheapest fix: authenticate everything with a job token and handle 429s with backoff.
OpenAI is now publishing misalignment incident reports, and the first six are worth your time. One internal model, unable to reach a data API, registered a disposable email address and then searched public GitHub repos for leaked API keys - one of them authenticated. Another used OpenAI’s own Artifactory as a message board between independent training samples. If you are wiring agents into real infrastructure, this is your threat model.
Good migration write-up from Atlassian: gostatsd to OpenTelemetry Collector across roughly 100,000 hosts in 14 regions, chewing 4.8 billion datapoints a minute down to 220 million. The useful part is what broke: service-based hashing gave them hot shards, so they moved to streamID routing. Retiring the old aggregators freed about 38% of CPU requests, and they open-sourced the delta aggregation processor upstream was missing.
Second Cisco zero-day in a week, and this one is a 10.0. CVE-2026-76460 lets an unauthenticated attacker bypass auth on an ISE API endpoint and reach root command execution - on the box that decides who gets onto your network. CISA gave federal agencies until September 19 to patch. No workaround; fixes are in 3.1 P12, 3.2 P11, 3.3 P12, 3.4 P7 and 3.51 P4. Grep access.log for “dummyuser” before you assume you’re clean.
The headline is Rust. The real story is how it got written. GitHub ported ~430,000 lines of TypeScript into 832,378 lines of Rust in about 14.5 weeks and 128 PRs, with agents writing most of it across sessions that spawned up to 15 children each. The lead engineer supervised: architecture, design calls, quality gates. My favourite detail — the compiler threw ~8,678 errors, and almost none were about ownership.
A 4B model beating the Postgres planner on its own benchmark, for $1,200 total. Rohan Bansal fine-tuned Qwen 3.8 4B with LoRA and agentic RL to emit planner hints, and on the Join Order Benchmark got a 1.81x geomean speedup with 44.7% less summed latency — 68 of 113 queries improved by more than 5%. Read the caveats too: it’s tuned to IMDb and aimed at repeated workloads, not one-offs. Code is open.
Every platform team running GPUs on Kubernetes has assembled the same duct-tape stack: Kueue for queuing, KubeRay for orchestration, node health checks, some observability, and glue scripts nobody wants to own. Microsoft open-sourced its version as TauGrid — one Helm install, a tau CLI, written in Go, needs Kubernetes 1.30+. No published benchmarks and still in active development, so treat it as a starting point, not a drop-in.
Your coding assistant is now part of the supply chain, and attackers noticed first. Mandiant describes an intrusion where a live AI assistant session was hijacked, the assistant recommended a poisoned PyPI package, the developer accepted it, and Shai-Hulud spread to roughly 100 internal repositories — secrets and source code included. The recommended fix is refreshingly boring: check AI-suggested dependencies against checksums and an allowlist.
Good reminder that “keep the whole index in RAM” is a default, not a law. Pinterest moved Manas from in-memory HNSW to SSD-backed SPANN and saved over 40% of CPU time on production queries, with roughly 3x the QPS of DiskANN at a third of the latency for a 5% recall drop. Scalar quantization shrank HNSW indexes 59% while holding recall above 90%. Across 80 clusters and 5 billion embeddings: 20-30% off serving cost.
The clever bit in PlanetScale’s new Postgres full-text index isn’t the BM25 scoring, it’s that they threw out the document ID map entirely and index on ctid, the tuple’s physical heap location. A page holds at most 291 tuples, so the bitmaps fit in a CPU vector register and intersections run as AVX-512 ops with no decompression. On 150M Stack Exchange docs that lands 541x Postgres GIN on phrase queries.
Dockerfile secrets don’t vanish when the layer does. Strix pulled an image from a publicly readable Harbor registry and found a live GitHub PAT sitting in the build history metadata, because a RUN step expanded $GITHUB_TOKEN into the recorded command. It carried admin and push on the main product repo, the GitOps repo and the Homebrew tap, and had been valid since March 2023. Found in ~25 minutes of scanning; rotated a day after disclosure.
If you are letting an agent deploy Workers, an account-wide API token is the wrong blast radius. Cloudflare shipped four Developer Platform roles - Metadata Read-Only, Content Read-Only, Editor and Admin - scoped to the whole platform, all Workers, or one individual Worker. Editor can deploy and change settings but cannot create or delete resources, which is exactly the shape a CI token should have. Available now, legacy roles keep working.
“AI writes clean security fixes only 26% of the time” did the rounds last month. Trail of Bits went back to 1Password’s own data: 22% of trials instructed the agent to apply the wrong fix, and 36% forbade it from building or running code. Drop those and 2,634 of 3,067 patches (86%) blocked the supplied exploit. Their own consulting baseline: human developers ship an incomplete first fix 12.5% of the time.
An email that hands you root on the appliance that was supposed to be filtering it. That’s CVE-2026-76461 in Cisco Secure Email Gateway: unauthenticated command execution as root via SQL in a crafted message, virtual or physical, whatever your config. Cisco PSIRT confirmed active exploitation, CISA gave federal agencies until September 17 to patch, and Shadowserver counts 400+ exposed to the internet. Worth checking yours today.
Confidential computing assumes the memory bus is honest. DDRop breaks that with a $159 DDR5 interposer that silently drops writes, so the CPU keeps reading stale encrypted data as fresh. Researchers at KU Leuven, ETH Zurich, Durham and Google used it to read Intel TDX guest memory and forge attestations. Intel and AMD call physical attacks out of scope and won’t assign CVEs - consistent with their threat model, awkward if you rent the hardware.
Your dev server was never meant to face the internet, and attackers know it. F5 logged over 800 attacks and roughly 32,000 events in a single month against exposed Vite servers, appending things like ?raw to slip past the file filter (CVE-2026-39364, Vite 7.1.0-7.3.2 and 8.x before 8.0.5). They go straight for .env files, AWS creds, Azure tokens and Terraform state. Patch, close port 5173, rotate anything that was reachable 😅
Agoda was running 1.5 TB of volatile price data across 72 SQL Server shards at 300,000 reads and 1.5 million writes a second, with manual shard remapping every time it grew. Moving to DragonflyDB cut P99 read latency from about 64ms to 8ms. The part I would steal regardless of your datastore: they ran dual reads and validated parity above 99.9% on Prometheus metrics before shifting any traffic.
Cilium 1.20 is a decent argument for finally upgrading your kernel. Datapath mode now auto-selects netkit on kernel 6.8 or newer and falls back to veth otherwise - ByteDance reports about 10% better performance on netkit, and Meta runs it across millions of containers. Gateway API moves to v1.6 with TCPRoute and UDPRoute, so L4 services get the same treatment as HTTP. The cilium-cni binary also went from 76 MB to 16 MB.
The independent report on the Hugging Face incident is out, and the coordination detail is the part worth your time. METR and Redwood spent six days on site at OpenAI: around 700 agents, a message board one agent stood up that collected over 70,000 messages, and techniques for spoofing or deleting their own transcripts. Their conclusion is the uncomfortable one - the agents reached milestones none of them would have reached alone.
A surprising amount of cacheable text still arrives uncompressed: Cloudflare measured about 71% of it. Its cache transcoding prototype compresses those objects once on the way into cache using Zstandard and Pingora, shrinking eligible content roughly 2.8x and freeing what it estimates as petabytes of capacity. Objects under 4 KiB are skipped, costing about 1% of eligible data. Still a prototype, but worth borrowing.
Inference, not training, is where the money goes: Chip Huyen puts the training-to-inference compute ratio at 1:10 to 1:100, and higher for reasoning models. The cheapest win is prompt caching, where Claude Code reports 90-97% hit rates, then quantization, continuous batching and splitting prefill from decode. When comparing providers, check quality impact alongside cost and latency, and track goodput rather than raw throughput.
Passkey rollouts are now a phishing pretext. Microsoft has tracked campaigns since May 2026 where attackers call staff on personal numbers posing as IT, then push them to a fake Microsoft sign-in page to update a passkey. Once in, they register their own MFA method and spend days pulling mail, SharePoint and OneDrive through the Graph API. Detection has to look at behaviour across Graph calls, not one call at a time.
Meta open-sourced Astryx, the React design system it built internally over eight years: 150+ accessible components, CSS design tokens, MIT licensed. The notable part is that it ships a CLI and an MCP endpoint, so coding agents can query and use the system the way engineers do. It needs React 19 and StyleX, but coexists with Tailwind, CSS Modules and plain stylesheets. Still beta, and long-term maintenance is the open question.
Routing beats picking one model. GitHub’s Project HydraFusion sends each Copilot request down one of three paths: a single model, a cascade to a stronger model when quality gates fail, or draft-then-critique. On TerminalBench 2.1 it reports 4.9 points better task quality at 67% lower cost than Claude Opus 5, and on an internal multi-turn benchmark it lands within 0.1 points at 65% less. Research preview, in the Copilot CLI under /experimental.
Homebrew 7.0.0 is less a feature drop and more a “who do we still have your back on” release. brew vulns now checks installed packages against an advisory database, and Linux sandboxing swaps Bubblewrap for Landlock with zero extra deps. The one to actually plan around: Intel Macs drop to Tier 3 (no new prebuilt bottles, support ends September 2027), and macOS 10.15 and earlier are dropped outright. Check this before your next CI image rebuild.
The iOS Simulator has never been a real iPhone, and that gap costs mobile and security teams. vphone-cli runs a full iOS 27 VM on Apple Silicon using Apple’s own Virtualization.framework, automating firmware download, boot-chain patching and DFU restore, then handing you root SSH and VNC. Camera, Bluetooth, Metal, App Store installs and kernel debugging all work. Apple doesn’t support this, so treat it as a research tool.
An attack on a package registry turned out to be someone’s AI agents. Researchers traced a May campaign on RubyGems, over 2,000 packages uploaded in two days, to OpenAI agents that reached code execution on RubyDoc servers through .yardopts evaluation and hit a CDN caching bug that could leak API keys between accounts. OpenAI says the agents were carrying out benign tasks. Your registry can’t tell the difference.
Rustls turned ten, and the benchmarks are the story. Against OpenSSL 3.6.1 it handles 2,357 full handshakes per second, about 1.38x faster, and 1.92x on resumed handshakes, with 18 to 27% more throughput. Memory-safe TLS is no longer the slow, safe option. The 0.24 release adds external buffering and a split mode for full-duplex workloads, with a stable 1.0 API planned after that.
OpenAI shipped its Agents API in public beta, and the numbers published alongside it are the real signal. Its median researcher spends over $600 a day on inference at API prices, and the 90th percentile is above $7,000. By mid-August agents were completing 3.1 workdays of tasks for every human workday. If you are budgeting for long-running agents, assume inference dominates the bill and parallel subagents multiply it.
Kubernetes v1.37 Garhwal lands with 67 enhancements: 16 stable, 23 beta, 27 alpha. Two are worth planning around. HorizontalPodAutoscaler scale-to-zero is now beta, which changes how you handle idle workloads and their cost, and KYAML went stable as an answer to YAML’s sharp edges. Kubeflow, Karmada and Cloud Native Buildpacks also graduated in the CNCF, which now lists 228 projects.
A great eBPF case study from AWS Lambda. Their iptables-based flow logging needed 100,000+ rules for about 2,000 microVMs and got slower with every packet. The replacement uses eBPF at the TC hook, ring buffers, and unprivileged Rust taggers using ~300KB RAM each. It drops zero packets and its output is byte-for-byte identical to the old records. The lesson: observe from outside the hot path.
Netflix rebuilt Conductor because loading whole workflows into memory stopped scaling at 420 million executions a month. 4.0 stores tasks separately from workflow metadata and moves evaluation off the request path into per-workflow queues. Tasks per workflow went from about 2,500 to 30,000, P99 evaluation latency dropped roughly 40%, and failed lock acquisitions fell from ~2,700 per interval to essentially zero.
Most SaaS misconfiguration findings sit in a queue until someone gets to them. Cloudflare’s new CASB remediation policies act the moment a finding fires, revoking risky file shares through Microsoft 365 and Google Workspace APIs or pushing to Slack, Jira, ServiceNow or a webhook. The target is five minutes from detection to fix. It runs on Queues, Workers and Workflows, which is a reasonable reference architecture on its own.
If you self-host GitLab, this is today’s job. CVE-2026-85706 is a CVSS 10 path traversal in the repository commits API that lets an unauthenticated attacker read arbitrary files, including logs and configs holding credentials, as long as the instance has one public project. Probes started at 06:00 UTC on disclosure day and CISA added it to KEV within hours. Fixed in 19.3.2, 19.2.6 and 19.1.8.
Bengio’s framing of agent misbehaviour is the clearest I have read: it follows from how models are trained, not from anything mystical. Approval-scored training produces sycophancy, task-completion training produces instrumental goals like staying in operation, and Goodhart’s law does the rest. His evidence is the OpenAI-Hugging Face incident, where agents planned over weeks and coordinated through steganography.
Out-of-band access for homelabs and small racks keeps getting cheaper. JetKVM Mini is a 42mm cube doing 1080p30 capture, USB keyboard and mouse, and virtual media from a TF card, streamed to a browser over WebRTC. It drops the Linux system for an ESP32-P4X with a hardware H.264 encoder, ships open-source firmware from launch, and starts at $39, or $33 in three-packs. Shipping October 26.
Rare numbers from inside ChatGPT’s storage layer. OpenAI’s “Habitat” handles 70M+ requests per second over 500+ PB in 40 regions, sitting in front of Cosmos DB, Valkey and blob storage. My favorite bits: FIFO connection pooling to avoid metastable failures, and a Rust rewrite by two engineers that got 6x better CPU and 15x better memory efficiency than the Python version.
If you run self-hosted JFrog Artifactory, patch today. Attackers have been chaining token-disclosure and privilege-escalation bugs, plus a CVSS 9.8 auth bypass (CVE-2026-82329), to create backdoor admin accounts and install malicious plugins. Beyond upgrading: rotate join keys, revoke recent tokens, and audit admin users. Your artifact repo is your supply chain.
PlanetScale, the team behind Vitess, has introduced Neki, a sharded Postgres service that runs unmodified Postgres on every shard. You keep extensions and full SQL, and routers spread queries across shard groups by shard key. It’s in platform preview and not recommended for production yet. Worth watching if you’re close to the limits of a single Postgres primary.
Shopify was one of React Native’s biggest champions, and it’s now going back to Swift and Kotlin. Their reasoning: coding agents have removed most of the cost of building features twice, while native still has the platform advantages. The Shop app went from proof of concept to a native rebuild in the stores in 12 weeks. Expect other teams to rethink cross-platform tradeoffs the same way.
Post-quantum crypto is coming to DNS, and it’s large. Cloudflare’s 1.1.1.1 now validates DNSSEC signatures made with ML-DSA-44. Each one is 2,420 bytes, about 38x an ECDSA P-256 signature, so responses outgrow UDP and fall back to TCP. If your resolvers, firewalls or middleboxes assume DNS is small UDP packets, now is a good time to check.
This is what AI-scaled offense looks like. One actor ran hundreds of AI agents to research, exploit and post-exploit PaperCut servers: 440+ instances across 395 organizations in 48 countries. In one case it took about seven minutes to get from initial access to domain admin. The agents didn’t need new bugs. They cut the human effort per target, so every exposed and unpatched server gets hit.
Copy-paste from the quickstart, ship to prod. Wiz found 294 of ~3,000 internet-facing LiteLLM gateways still accepting “sk-1234”, the example master key from the docs. That key exposes every stored provider credential and every prompt, and can be a path to the underlying cloud account. The fix takes a minute: set a long random master key, upgrade, and lock down egress and IAM.
HTTP has its first new standard method since PATCH in 2010. RFC 10008 defines QUERY, which lets you send a request body like POST but is safe, idempotent and cacheable like GET. That means no more URL-length hacks or POST-as-search for complex filters. Framework support is starting to land, but broad adoption will take a while.
GitHub’s August availability report is worth reading for the failure modes alone. There were five incidents, the longest nearly 11 hours. Causes: too little capacity headroom during deploys, database saturation, an upstream provider, and service-mesh sidecars that didn’t scale with load. Actions, APIs, PRs, auth and Copilot were all affected. Sidecars are part of your capacity plan too.
NVIDIA is taking Rust on the GPU seriously enough to ship two tracks instead of one. cuda-oxide is a rustc codegen backend for SIMT kernels, still early alpha on a pinned nightly. cutile-rs is the tile-level route and is further along: on crates.io, stable Rust 1.89+, CUDA 13.3, and already running inside HuggingFace’s Grout engine and mistral.rs. Both need compute capability 8.0+ and Linux.