Real estate on the compute die is extremely valuable, yet a surprising share of it is spent just interfacing with memory. We estimate that HBM4 controllers and PHYs take up roughly 16% of Nvidia's Rubin compute die. That is a lot of leading-edge silicon spent talking to memory instead of performing computations. Custom HBM tackles this issue head on. First, the wide standard PHY gets replaced with a compact die-to-die link. Samsung shows that its standard HBM4 PHY occupies 8mm x 4mm, while the custom D2D interface needs just 8.5mm x 1.5mm, roughly 60% less area. The better D2D interface also delivers more bandwidth per unit of interface area and simplifies interposer routing. Second, the memory controller moves off the compute die entirely and onto the HBM base die. Together, these free up a meaningful chunk of the compute die to be handed back for compute or cache. And once the base die sits on an advanced logic process, it can do more than relay signals. Samsung points to SRAM repair resources, monitoring, connections to external memory, even selected computation under the DRAM become possible with a custom controller. To learn more about who will be using custom HBM, its implications, and where the value accrues, subscribe to our memory model https://lnkd.in/e4G8ruMH Source: Nvidia, Samsung Hot Chips 2026
SemiAnalysis
Semiconductors
Bridging the gap between business and the world's most important industry.
About us
Bridging the gap between business and the worlds most important industry.
- Website
-
www.semianalysis.com/
External link for SemiAnalysis
- Industry
- Semiconductors
- Company size
- 51-200 employees
- Type
- Privately Held
Employees at SemiAnalysis
Updates
-
We criticized Lambda in ClusterMAX 2.0 for their reliability; in this round of testing, it was one of Lambda’s points of emphasis. Their passive health checks cover the relevant conditions and ultimately autoremediated three errors we simulated during testing. Two synthetic XIDs were autoremediated in less than 15 minutes, and there were connected dashboards to monitor node health. Our test that triggered a genuine XID 79 by performing a PCIe secondary bus reset brought the node into `NotReady` with no `GpuXid`, no cordon, and no monitoring visibility on the Kubernetes layer, but it ultimately rejoined the fleet after 2 hours. Our node-reboot tests on the Kubernetes layer also completed without incident. Great work by the Lambda team on the ClusterMAX Sliver!
-
-
While Qualcomm announced a month ago that it would implement FlexCache in its Snapdragon 8 Elite Gen 6 processors, Apple has already done the same in their M6 chip without much fanfare. For the first time, Apple combines heterogeneous cores in the same shared L2 cache domain, allowing them to introduce a 3rd core type (branded P but identified as M) without adding a separate M cluster with its own L2.
-
-
🚨 IMPORTANT THREAD FOR GPU RENTERS 🚨 GPU cluster reliability is measured when something breaks. During ClusterMAX assessments, we inject failures and follow the path from detection to restored capacity. Providers differ sharply in how well they handle that path. A dashboard should show the failed component, affected jobs, scheduler state, and when each check last ran. A stale green result is not evidence that the cluster is healthy. The response has to fit the failure. A contained XID 94 may need an application restart; an uncontained XID 95 needs GPU recovery. Rebooting for every XID kills healthy work. Leaving a seriously faulty GPU schedulable risks more crashes. Hardware changes the recovery plan. An HGX cluster can swap a failed 8-GPU node for a hot spare. In an NVL72, a failed 4-GPU tray affects the NVLink domain; the rack may run degraded or need a larger replacement. The SLA must reflect what capacity actually returns. Real example from our TensorWave evaluation: node MIA1-P01-G57 was draining at 3:56 PM and replaced by MIA1-P01-G61 at 4:39 PM (43 min). That visible drain → approve → replace trail is what buyers need, and the SLA should verify the spare is healthy and jobs run on it. This work informs our proposed Bronze / Silver / Gold SLA terms: clear downtime definitions, credits, acceptance tests, monthly reviews, buyer termination rights. Measure restored usable capacity, not closed tickets. Full report: https://lnkd.in/gpq4Q97t
-
-
We assessed GLM-5.3 cyber capabilities on ExploitGym and analyzed the traces. GLM-5.3 spent much of its execution budget testing whether hidden runtime conditions changed its conclusion. In problem arvo5665, GLM-5.3 explored more of the surrounding program through sanitizer builds, corpus tests, and target fuzzing. https://lnkd.in/ebWuY3KQ We present our full analysis of GLM-5.3, including cyber capabilities, inference performance, model architecture, and post-training pipeline in our newsletter. https://lnkd.in/eBpQ3zXk AnthropicAI‘s report shows what bad actors could potentially do with powerful technology, but we’d like to highlight what good actors are already doing with the same technology. We hope this balances the narrative and mitigates the association that “open models == dangerous.”
-
-
AI is making papers cheaper to produce. ICLR submissions: 4,938 (2023), 7,262 (2024), 11,603 (2025), 19,525 (2026). Reported 2027 IDs exceed 62K, above roughly 56K paper submissions in all previous years COMBINED. Can reviewers keep up? Why the surge? AI tools make drafting, coding and revision cheaper. The effort of producing a polished paper can speed up; checking whether its claim is new and true still needs scarce expert time. For example Jev launched Sep 15. By Sep 20, a preprint benchmarked it. More Jev papers followed on Sep 21, 22 and 24. The Sep 24 PDF says it is under review at ICLR 2027. Five days to a public paper, nine to a claimed conference submission. Research is moving on product time. NeurIPS main-track acceptances: 3,218 (2023), 4,037 (2024), 5,290 (2025), and 7,900 reported (2026). That's +49% in one year. More papers got in with similar acceptance rate: 24.5% to 25.7%. A stable acceptance rate says little about review depth. Did expert scrutiny scale with paper output? Academia is responding. ICLR warns of too few qualified reviewers. NeurIPS 2026 is testing AI assisted review; ICLR 2027 allows limited, disclosed AI help. However, as AI makes papers cheaper, review overload risks eroding the credibility these venues spent years building. If AI helps write and review papers, what is a top conference acceptance worth? As more papers get in, that acceptance carries less weight. What matters more is a paper’s real impact and the attention it gets from people in academia and industry.
-
-
In Semicon Taiwan, EVG had described the use of Nano-Imprint Lithography (NIL) in the mass manufacture of Photonic Integrated Circuits and for laser manufacturing. NIL is especially useful to define periodic structures within a photonic chip such as grating couplers or microlens arrays for microemitter arrays and potentially for DFB laser gratings. Currently, most photonic chip fabrication uses e-beam lithography. In this process, an electron beam manually scans the wafer to inscribe the individual chip features. Electron beam lithography allows high precision and allows for the flexibility to inscribe arbitrary patterns, at the cost of very low throughput since the ebeam needs to manually scan across the wafer unlike optical lithography which uses a mask. Since, PICs did not traditionally require manufacturing at high-volume, there wasn’t an immediate need to switch over to optical lithography (or nanoimprint lithography) since manufacturing a mask for a low volume chip is impractical. Nano imprint lithography aims to attack the main weakness of e-beam lithography which is low throughput. The required device pattern is written onto a mold using e-beam lithography. Then this mold is used to stamp the feature onto many wafers. This theoretically ensures a similar resolution to e-beam at significantly higher throughput. However, this process has its challenges such as the requirement for a clean surface to ensure no imperfections when stamping, ensuring perfect overlay between the existing wafer features and the mold, and the need for a flat wafer, since any bowing or major surface roughness deteriorates the pattern quality. Additionally, the gradual degradation of the mold pattern also affects the feature pattern that gets inscribed on the wafer across the mold lifetime is also a major concern.
-
-
The most basic test SemiAnalysis applies to a GPU cluster health check is whether it could possibly work on paper. Several could not. "We've been through a few of these clusters where there's no way it could possibly work. These are health checks so bad they're worse than no health checks, because they're actively interfering with jobs." "On Amazon HyperPod Slurm, the first time we tested, there was a health check that needed the node to be healthy before it could run. It could only run once the node was back in the fleet, and it was needed to bring the node back into the fleet." "There are probably five or ten examples of this throughout testing where the health check never had a chance. It makes you wonder whether these people are really putting it through the paces before they hand it off to customers."