
Why Do Mainframes Still Exist? What's Inside One? 40TB, 200+ Cores, AI, and more!
What this covers
Dave Plummer explores the IBM z16 mainframe from design to assembly and testing. What's inside a modern IBM z16 mainframe that makes it relevant today? For info on my book on ASD/Asperger's, please check out: https://amzn.to/47ItFnR
0:00 Introduction 4:47 Inside the z16 7:03 Super Input Output 9:39 Factory Assembly 10:56 Accelerators 12:24 Test Lab 13:50 DIMM installation 14:42 Water Cooling 16:49 Why Mainframes? 18:20 Fiber 19:47 Conclusions
Errata: Around 5:30, the CPUs are LGA(Land Grid Array), not PGA At about 7:07 should be 200 cores, not 256. At 20:03, they now achieve eight nines of reliability for both z16 and its Linux-only counterpart, the IBM LinuxONE 4. In a LinuxONE config, max RAM is 48TB.
Source description (no synthesized summary yet).
IBM mainframes remain relevant and superior to distributed commodity systems for mission-critical workloads because their vertical integration, reliability engineering, and channel-based I/O deliver unmatched performance, availability, and cost efficiency at scale.
- Mainframes achieve seven-nines availability (milliseconds downtime annually) through redundancy, hot-swappable components, and deterministic performance that commodity systems cannot match
- Financial and retail transactions require immediate, guaranteed performance that mainframe architecture delivers but distributed systems cannot guarantee reliably
- Mainframe consolidation reduces total cost of ownership through lower administration, energy, software licensing, and datacenter floor space costs despite higher upfront investment
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
Z16 mainframes can encrypt main memory using quantum-safe algorithms, so even if a memory page is captured today, a quantum computer still cannot crack it 50 years from now
“main frames like the Z16 can even encrypt main memory and they do it with Quantum safe algorithms so that even if you took a snapshot of a memory page today a quantum computer still won't be able to crack it 50 years from now”
The Z16 CPU achieves a base clock and boost clock of 5.2 GHz simultaneously without throttling or dynamic clock adjustment; when workload reduction is needed, the system inserts wait states rather than lowering frequency, allowing consistent performance across all cores under full load.
“each CPU runs at its full 5.2 GHz that's the base clock and the Boost clock and it doesn't change they never dclock or throttle this CPU so even if they need to slow things down they simply insert weight States rather than changing the frequency”
Mainframe consolidation reduces total cost of ownership by enabling management of one mainframe instead of hundreds of servers, leading to savings in administration, energy consumption, software licensing, and datacenter floor space
“mainframes consolidate workloads meaning that instead of managing hundreds of servers a business might only need to manage one Mainframe this consolidation leads to Savings in administration energy consumption and even software Licensing in many cases it might not be sexy but conserving floor space in the data center can translate into serious savings over time”
Real-time credit card fraud analysis on every single transaction requires AI inferencing speed and reliability that only mainframes with on-chip AI can deliver, whereas generic transaction processing without true real-time AI fraud analysis would be insufficient
“if you're a credit card processor without something like this you simply would not be able to do true AI based fraud analysis on every single transaction but with a Z16 you can it's one thing for chat GPT to snooze for a few seconds but certain transactions need to be smart and fast”
Mainframes have robust security features ensuring data protection and minimization of security breaches, with only about one-tenth of 1% of mainframe customers experiencing breaches compared to significantly higher rates on commodity systems
“main frames come equipped with robust security features ensuring that data remains protected and that system disruptions due to security breaches are minimized only about one/ tenth of 1% of Mainframe customers ever experience a breach a number that can be significantly higher on commodity systems”
Z16 systems can contain a mix of virtualization, bare metal, and multiple operating systems and environments running side by side, each completely isolated from one another, with each VM or environment believing it has complete control of the system.
“not only can you partition the system into numerous VMS but those VMS can also contain other hypervisors you can have a mix of virtualization bare metal and multiple operating systems and environments all running side by side at once completely isolated from one another with each believing that it has complete control of the system”
Dave concludes that even in an age of rapid technological advancement, mainframes have retained their relevance and remain a cornerstone of IT infrastructure for large organizations due to unmatched reliability, ability to handle vast data efficiently, and long-term cost savings.
“even in an age of Rapid technological advancement main frames have retained their relevance their unmatched reli ability combined with their ability to efficiently handle vast amounts of data and offer cost Savings in the long run ensures they remain a Cornerstone in the it infrastructure of many large organizations”
Mainframe channel-based I/O systems deliver higher throughput and more reliable I/O operations than bus-based systems in commodity PCs, functioning like dedicated highways for data whereas PCs share a common bus like traffic jams
“main frames use a unique channel-based iio system unlike the bus based systems in most PCS and this channel-based design allows for higher throughput unlike the bus based systems in PCS this channel-based design allows for higher throughput and more reliable iio operations it's a kin to having dedicated highways for data ensuring smooth and efficient traffic flow imagine that everywhere you wanted to drive there was a fresh paved road with no other traffic on it and then contrast that with sitting on a bus in a traffic jam”
IBM uses robotic automation to insert memory modules into Z16 mainframes with extreme precision, virtually eliminating human error and allowing one operator to manage four stations running at approximately the same speed as a human operator
“the robot knows and so when it needs one it reaches over grabs one of the memory modules takes it over to a scanner which presumably both verifies it's the right one and Records where it's going the modules is in aligned much more precisely than a human could do...I'm told the automation of this particular task pretty much eliminates errors and mistakes and it allows one operator to run four stations that run at approximately the same speed as a human”
IBM mainframes undergo extensive validation testing including vibration and shock testing where g-loads experienced on delivery routes are recorded and then played back condensed in time on hydraulic shock tables to simulate real-world delivery stress
“IBM subjects its designs as well as individual systems to some impressive testing they do extensive vibration and shock testing wherein they record the GS experienced on an extended delivery ride then they play back that g-load information condensed for time on a hydraulic shock table”
Financial services use cases are concentrated in mainframe domains, including ATM cash withdrawals, account balance checking, retail purchases, airline ticketing, hotel reservations, tax filing, and package delivery tracking, all of which run on mainframes behind the scenes
“you can find those cases con concentrated rather heavily in the financial sector if you've ever withdrawn cash from an ATM odds are it was a main frame behind the scenes if you've checked your account balances from your phone that's a Mainframe if you bought a laptop at the Apple Store or bought an airline ticket or booked a hotel room it happened on a Mainframe somewhere and if you filed your federal taxes or had a package delivered by UPS or FedEx again it's a Mainframe”
Z16 mainframe assembly uses just-in-time parts delivery with carts arriving at assembly stations bearing order numbers and the specific parts required at the next assembly station, similar to bespoke automobile manufacturing
“carts arrive here in the assembly area with numbers that correspond to customer orders and that include parts required at the next assembly station”
Drawer interconnects on Z16 mainframes operate at approximately 320 GB per second, which is about 10 times faster than the ~25 GB per second that a modern Xeon processor can achieve when accessing main memory at 3200 megatransfers per second
“each can communicate continually at the maximum rate of the connection that's also the speed at which one CPU can access the cache of another CPU in The L4 layer and it looks to be 320 GB per second for comparison a modern Zeon running memory at 3200 megat transfers per second of 64-bit Ram approaches about 25 GB per second”
Z16 memory can be over-provisioned, meaning a system might have 40 terabytes installed but only activate 30 terabytes, with the remainder available for activation as needed and system upgrades without reboot
“memory can be over-provisioned meaning you might install 40 terabytes but really only activate 30 terab of it the remainder then can be activated as needed and the system can be upgraded without a reboot”
Z16 mainframes implement RAIM (Redundant Array of Independent Memory) in RAID 10 fashion, where data is striped across RAM to protect against cosmic ray bit flips that ECC memory alone cannot correct
“Z16 does r where M is for memory it effectively Stripes data across Ram in RAID 10 fashion meaning that if your system is hit by a cosmic aray that flips bits in a way that ECC memory alone could never hope to fix nothing is impacted”
Each CPU drawer in a Z16 is connected directly to every other drawer so they don't share bandwidth over a common bus, and each can communicate at maximum connection rate continuously
“each of these drawers is connected directly to each of the the other drawers so they don't have to share bandwidth over some common bus and each can communicate continually at the maximum rate of the connection”
Even with unlimited computing resources available, distributed commodity systems cannot match mainframes for certain workload classes because not all large databases can be effectively distributed and some tasks inherently require centralized architecture
“it still seemed to me that if you gave me pretty much any task that a Mainframe was designed to do I was confident that I could think of a way to do it with a rack full of comparatively cheap PCS and ssds after all for any problem you can just throw more boxes at the distributed PC version and they're cheap and that's all true for the most basic Cloud tasks like web serving in reality however not everything is as easily distributed as web serving there still exist large Uncharted databases that cannot be effectively distributed for example in fact there are numerous cases where the main frame still shines as the best solution”
IBM invented virtualization in the 1960s for mainframes, allowing division of a single mainframe into multiple virtual systems each running its own operating system, long before virtualization became common in commodity systems
“IBM has been doing virtualization since long before virtualization was cool going back to the 1960s”
Z16 mainframes support various specialized PCIe cards including storage cards, FICON high-speed data transfer boards, cryptography assist boards, compression boards, and AI inference boards
“what do Mainframe customers actually run for PCI cards well storage cards of course but also things like ficon high-speed data transfer board cryptography assist boards compression boards AI boards and so on”
Z16 mainframes use closed-loop water cooling with deionized water, manifold distribution to each CPU socket, and hot-swappable CPUs that can be replaced without draining fluids from the system
“all of the higher spec main frames run a water cooling Loop in a closed circuit the lines enter the back and then join a log style manifold which has both a Inlet and a return to each CPU socket...You can hot swap a CPU by simply removing the mounting bracket and then popping the CPU Chiller off the top of the CPU itself you then swap the CPU put the bracket back and there's no need to drain any fluids from the system”
IBM Z16 mainframe CPUs contain two dies with eight cores each, totaling 16 cores per chip, with each core having its own private 32 megabytes of cache, which is significantly larger than the 64-96 kilobytes per core found on typical server CPUs
“each module features two CPU dieses and each die contains eight cores for a total of 16 per chip...each core has its own private 32 megabytes of cache which completely dwarf C 64 to 96 kilobytes per core that you'd find on a typical server CPU”
Z16 mainframes are available in four main configurations: a giant 4-cage monster with ~200+ cores, a single-frame mini monster with ~64 cores, and rack-mount versions with 18u to 42u space requirements
“generally you're going to buy a main frame in one of four configurations either the giant 4 cage monster with 200 and some cores or a singleframe mini monster with about 64 cores and since everything is built on a 19-in rack standard they also Supply rack mount versions with as little as 18u and as much as 42u requirements for space”
IBM datacenter fiber infrastructure uses multiplexing technology to merge a dozen fiber optic signals onto a single output using color filters, and maintains a 100 km fiber loop to physically distance machines and test replication and latency over true distances
“this box can Multiplex a dozen fiber optic signals onto a single output by merging them through color filters and then multiplexing them all onto the same glass line they also have a 100 km fiber Loop in order to physically distance two machines from one another and test things with forc latency so that replication and so on can be tested over true distances while you're in the same room”
Mainframes achieve up to seven nines of availability, meaning just milliseconds of downtime annually, which is crucial for mission-critical applications where even brief outages have significant consequences
“main frames especially models like the IBM Z16 are renowned for their reliability they can achieve up to seven nines of availability in Practical terms this means just milliseconds of downtime annually such reliability is crucial for Mission critical applications where even a brief outage can have significant consequences”
Mainframe virtualization allows partition isolation where if one partition encounters an issue it won't necessarily affect others, ensuring continuous operation even when individual partitions fail
“this functionality allows a Mainframe to divided up into multiple virtual systems each one running its own operating system that means if one partition encounters an issue it won't necessarily affect the others ensuring continuous operation”
IBM's mainframe cooling fill station still runs Windows XP, despite Windows XP being legacy software, because it prioritizes absolute reliability in the fill station over using modern software
“here in the center of this three-frame stack you can see the cooling Reservoir box...here's where I discovered they were relying on an old friend something I'd worked on well not the embedded version but Windows XP at least that's right their fill station still runs Windows XP I guess when you absolutely positively need your fill station to work every time you only rely on the very best”
Dave joked about whether 192 PCIe slots in Z16 could be filled with 192 GPUs for Bitcoin mining; IBM engineers appeared to disapprove, citing heat and power concerns, but noted that IBM would not create 192 slots if nobody was actively using them.
“I knew what everybody's first question would be so I asked whether you could just install 192 gpus and do some serious Bitcoin mining they seem to frown on that idea for a number of reasons not the least of which is probably heat and power concerns but rest assured they would not make a system with 192 slots if somebody somewhere wasn't actively using them”
Dave initially believed he could solve any mainframe workload by distributing it across cheap commodity PC systems with solid-state drives, but through visiting IBM reconsidered this belief upon recognizing fundamental workload classes that cannot be distributed
“at this point here's the dilemma I was facing even though I now knew a lot more about the systems it still seemed to me that if you gave me pretty much any task that a Mainframe was designed to do I was confident that I could think of a way to do it with a rack full of comparatively cheap PCS and ssds after all for any problem you can just throw more boxes at the distributed PC version and they're cheap and that's all true for the most basic Cloud tasks like web serving in reality however not everything is as easily distributed as web serving”
Z16 systems use preformed thermal transfer sheets rather than thermal compound for the interface between the heat sink and chip lid; this provides more consistent thermal properties and reliability compared to paste-based approaches.
“they don't use thermal compound but instead pre-formed sheets of thermal transfer material”
IBM sources fiber cables by the forklift pallet due to the enormous quantities needed for datacenter fiber infrastructure
“and when you use that much fiber how do you buy It Well by the forklift pallet”
Z16 mainframe CPU modules use a Samsung 7 nanometer process technology for chip fabrication.
“even using a modern Samsung 7 nmet process the telm cpu's architecture allows it to share unused L2 cache”
Dave had personal experience working at IBM summers and evenings during his college years, and his wife Nicole also worked at the IBM data center during college, and her father spent more than 30 years working for IBM
“what you may not know is that before that I actually put myself through school in part by working Summers and evenings at IBM my then girlfriend and now wife Nicole also worked at the IBM data center during college and her dad spent more than 30 years working for big blue”
Dave worked at Microsoft in the 1990s when the company culture was characterized by enthusiasm and belief that Microsoft was changing the world while making people rich; he observed IBM engineers had similar levels of genuine enthusiasm and passion for their work, something he has not witnessed firsthand in the tech industry in quite a while.
“I worked to Microsoft back in the 90s when everybody thought they were changing the world and getting rich at the same time so of course they were excited back then but the folks that I spoke with at IBM had the same level of enthusiasm and it's something I haven't seen firsthand in the industry in quite a while”
Dave visited IBM engineers Matias Dian and Mike who each had about 20 years of working experience in CPU caching and core design
“we sat down with Matias Dian and Mike each of whom had about 20 years working on their individual area of expertise and things like CPU caching and core design”
Z16 main frame assembly uses lifting stations that can not only lift the entire mainframe but can also drop it up to 18 inches to bring upper bays down to more reachable levels for assemblers
“this lift station is cool and that not only can It lift the entire main frame in order to make the lower Bays more easily accessible to the assembler but because we're on a raised floor it can also drop the main frame up to 18 in to bring the upper base down to a more reachable level”
IBM provides a visible clear Z16 frame model with cutaway panels and transparent windows to show the internal organization of compute drawers, I/O cages, and power supply units.
“let's have a look at a clear Z inside a clear rack so we can get an orientation for where everything goes and what makes up the main frame itself here we have the compute drawers on top below that we see all the io cages below that”
During manufacturing testing, if shock tests caused systems to tip over, IBM engineers did not appear concerned or surprised, suggesting such events are either expected or part of the deliberate test protocol.
“is okay to be honest I think something went wrong with the hold Downs on that last one at least I hope they don't really tip them over like that very often do”