7/19/2026 at 9:44:56 PM
“It runs quite slowly, (486-levels of performance on a Ryzen 5000 series according to this discord channel)”I’m sure there’s a joke about that being a huge leap in performance for Windows on IA-64 over the past 23 years waiting to be made from this.
by qubex
7/19/2026 at 10:37:49 PM
To be fair, most of ia64 performance issues were in the more complicated compilers. The chip itself had potential, but essentially abandoned the DOS/Windows market inertia. =3by Joel_Mckay
7/19/2026 at 11:14:32 PM
No (sorry, I post this every time the Mythical Compiler Myth reappears), the problem with VLIW is that it fundamentally doesn’t work for anything with unpredictable memory access patterns, and modern general purpose computing has moved almost exclusively in this direction. For VLIW to work, you need to either guess correctly what is in cache or not have a cache at all; as soon as you mispredict what has been loaded, you stall while an OoO processor keeps going and a speculative OoO processor even keeps guessing. With multiple workloads on the same hardware (virtualization, multitasking, multitenancy) this becomes an intractable problem even in the presence of the magic compiler which can solve for software based unpredictability (branch likelihood and pointer chasing), because every context switch clobbers an unknown set of cache lines and blows the entire thing up.VLIW works for single workloads. It works exceptionally well for single workloads with no or explicit cache like DSP. You can trade the footprint and complexity from OoO for a wider execution unit and more SRAM. It works well for HPC, too, for the same reason. But for anything where more than one process exists, it just really doesn’t work, and that’s most modern workload.
Itanium also has a unique set of self inflicted issues due in large part to Intel trying to make a wide variety of cross compatible parts, but IMO even if they’d got it right, it still would have died.
by bri3d
7/20/2026 at 3:15:08 AM
Itanium had all sorts of specialized machinery so that you could issue a load that's required in the future, and then keep chugging away at the instruction stream, similar to what an OoO processor does with it's reordering logic.One example was "advanced loads" which allowed you to issue a load as soon as you knew the address, even in the face of potential pointer aliasing in the future, and then later complete the load when you actually hit a data dependency that requires it. https://devblogs.microsoft.com/oldnewthing/20150805-00/?p=91...
Another example is "speculative loads" which lets you issue a load before you even know if it's valid, such as unrolling a loop for an array that you don't know if is a multiple of the unrolled loop chunk length (and therefore might trap with a page fault if fully resolved). https://devblogs.microsoft.com/oldnewthing/20150804-00/?p=91...
by monocasa
7/20/2026 at 4:39:35 AM
These worked well for “known” intra-task pointer aliasing situations, but if you don’t know what will be in cache due to preemption of any kind, you still don’t know how many cycles the speculative loads will take, so you get the same stall risks across a dependency hazard.by bri3d
7/20/2026 at 5:32:21 AM
OoO cores have load stalls too after preemptions. Even something like a Apple M core basically has to stall if the load has to go out to the memory controller. The goal is to move far enough ahead in the stream that you can at least issue the next load, and advaned loads lets you do that. Some Intel cores will occasionally speculatively predict zero for a load, but AFAIK that got turned off as part of the Spectre mitigations.by monocasa
7/20/2026 at 8:23:04 AM
The difference is that this is not the programmer's responsibility in modern machines: instead we bake-in some hardware that watches the online state of the machine and then actively decides on what to do.If you're trying to create a compiler that approaches the effectiveness of this statically, you're condemned to do a ton of extra work (you're essentially writing an emulator for your CPU core, and then a compiler that uses that model to produce optimal code - even then, you might not be accounting for nondeterminism on the actual target machine, and some information is simply not accessible to you when you are not on the target machine)
by eigenform
7/20/2026 at 9:02:47 AM
Once again, Itanium isn't a normal, simple in order core that stalls on loads. It has a big table called the ALAT to allow you to start loads as soon as you know the address, and then finalize them later when you're out of other work to do.It doesn't require determinism to work. And was designed by people that were quite aware of what the instruction stream looks like to an OoO core as it issues out of the rob.
Itanium had other sins than "magic compiler" woes, or even the inherent unpredictability of memory accesses. Mostly that it, like Cell, and Netburst was designed for a world where dennard scaling didn't end like a brick wall. As well as internal politics of Intel making it so that they were a bit loose and fast with die area.
by monocasa
7/20/2026 at 6:20:18 PM
Yes, I'm pointing out the fact that the programmer is expected to manage the ALAT, and that these machines do not automatically recover from cases where your advanced loads are incorrect. That process is expected to be part of the instruction stream, and [we have collectively learned that] that's an expensive feedback loop.Yes, it doesn't require determinism to work, but it means that your performance is especially sensitive to nondeterminism because the compiler cannot account for loads and stores that occur online. The SDM (see vol 1, section 9.5.1) describes this idea pretty well.
by eigenform
7/21/2026 at 10:11:07 PM
You didn't generally need to manage the ALAT as a programmer.You can, in the same way that you can manually manage reservation station port residency as a programmer in an OoO core to get maximum throughput, but neither of those are required in normal programming.
by monocasa
7/20/2026 at 7:23:14 AM
VLIW is being heavily used in some AI inference tasks (by Apple, AMD, Google, mainly in optimizing latency). It also is a nice way to do inference on edge devices that aren’t using constantly advancing models. It’s pretty much dead in general purpose computing though, and lacks the versatility of GPUs to go beyond inference.by seanmcdirmid
7/20/2026 at 12:01:05 AM
We agree it probably would have still failed, but mostly it was the legacy code-motion compatibility/performance issues that were practically inescapable without refactoring millions of lines of code.gcc maintained the ia64 target a long time for unclear reasons, but it was also still inefficient on other platforms. The FOSS compiler worked, but that was its only performance metric that counted for many users. =3
Not sure why you think the Intel compilers were a myth, as they are still around working far better than gcc in many use-cases:
"An Overview of the Intel® IA-64 Compiler"
https://webdocs.cs.ualberta.ca/~amaral/courses/605/papers/In...
by Joel_Mckay
7/20/2026 at 4:40:38 AM
The myth I was referring to was the overarching theme that “VLIW would have been practical if only a better compiler existed,” which I don’t believe to be true for modern or Itanium-contemporary general purpose computing patterns.by bri3d
7/20/2026 at 5:01:24 AM
Indeed, people are still bad a parallelism today, and most compilers still suck at reliably unrolling abstracted concurrent source intended functionality.Very few modern languages handle parallel scaling gracefully, and bodged on CUDA still isn't great either. As Moore's law ends, people have to reevaluate how they approach traditionally monolithic architectural design. =3
by Joel_Mckay
7/20/2026 at 9:36:51 AM
Honestly, you're conflating two things.1. VLIW exposes microarchitectural details, locking them in like an ABI. Updates to the microrachitecture will require changes to the ISA, thereby breaking backwards compatibility with every generation.
2. The dominant programming paradigm is sequential code, often just old C code with heavy pointer aliasing and little compile time extractable parallelism. This reduces static parallelism in the code base and shifts it into a runtime problem. The CPU discovers the dependencies at runtime instead. If you built a programming language that exposes more parallelism (think something like ParaSail), this problem wouldn't be as big as it is for C style programming languages.
Nobody will build a language for architectures that suffer from backwards compatibility problems, so why bother? VLIW is primarily suited for ASIPs and not much else.
by imtringued
7/19/2026 at 11:02:39 PM
Yeah EPIC pushed too much effort onto compilers and removed the context awareness of out of order reordering etc and the result was just a lame duck.by qubex
7/20/2026 at 12:24:29 AM
[flagged]by Joel_Mckay
7/20/2026 at 1:15:07 AM
You've mentioned Intel's compilers a few times but they're not a magic thing that could make Itanic not-suck. Suppose they make general purpose code that's twice as fast as GCC when the whole computer is dedicated to running one specific process. They don't, but let's pretend. Even in that ideal scenario, it falls apart as soon as you're processing unpredictable input. VLIW sucks hard at chewing through data that's not precisely what it's expecting. You could probably make a video codec that performed really well, but it would be impossible to make a database that performed consistently not-badly. It's not that the existing compilers weren't good enough, not even the magic Intel ones. It's that it's not possible to pre-generate VLIW opcodes that do well when the input value isn't a homogeneous stream. Even if Intel's code was 2x better than GCC's — and again, it wasn't — twice abominable was still abominable.I wasted more time SSHed into an Itanic than most people had to, so I'm speaking from first-hand experience. Some very, very specific Itanium code was a bit faster than the equivalent x86 or similar. That was always some very tightly scoped thing that did the exact same tight loop a gazillion times, like en-/decrypting a data stream. Anything more heterogenous ran poorly, as in multiple times the wall-clock time of the same workload on the x86 server next to it that cost a tenth as much. And that's even when using the magic Intel compilers.
Itanium was bad. The compiler tech didn't exist to make it not-bad, and in retrospect I think it become obvious that it couldn't exist. It wasn't a matter of the compiler authors needing to be clever. It was more like making Itanium live up to the hype required making P=NP.
by kstrauser
7/20/2026 at 2:38:01 AM
My point was most modern software developers still do not understand systems at an architectural level. I don't buy the VLIW paradigm was inherently inferior argument, but the code written for it with mystery failure modes at the time was a mess.And ia64 was only around 23% slower and several times more costly than amd64 options at the time. Both toads had their warts, but one was actually usable by mere mortals.
Its true just about everything either ran slow on not at all on ia64... Thankfully, we were still running Sparc based platforms at that time, as the Intel clown-show still looked like more work. But I do empathize with the ia64 trauma, as the roll out had a lot of collateral damage in some firms as people started jumping ship early to avoid accountability. =3
by Joel_Mckay
7/20/2026 at 8:18:17 AM
One immediate advantage that rescheduling-capable cores have that rigid compiler-based ones don’t nor ever would is that reschedules can notice patterns in execution in the current context and keep the ALU fed from well-provisioned caches even going as far as speculative execution. IA-64 and typeless or typed-at-runtime (JIT scripts, for example, like JavaScript or Python) would’ve been implemented very differently on Itanium.by qubex
7/20/2026 at 4:13:12 AM
I think they meant EPIC https://en.wikipedia.org/wiki/Explicitly_parallel_instructio..., not EPYC https://en.wikipedia.org/wiki/Epyc.by abbeyj
7/20/2026 at 4:37:16 AM
Indeed, my point was not all chips do well with Desktop application loads.People were bad at handling parallelism with ia64, and still have problems today on better amd platforms with less janky compilers.
The fact an $800 chip still can beat a $14000 chip at some tasks probably should tell people something about concurrency scaling overhead. =3
by Joel_Mckay
7/20/2026 at 8:14:14 AM
EPIC was Explicit Parallel Instruction Computing, the underlying engineering architecture of AI-64. Epyc is an AMD brand name.And pivoting to concurrency on a rescheduler-intensive core as opposed to concurrency on a scheduler-based core isn’t very pertinent, particularly since we have 25 years of Moore’s Law between then and now.
by qubex
7/20/2026 at 9:01:19 AM
"AI" responses are silly, and optimal 24 core count efficiency premise in Desktop applications offer diminishing returns on a highly concurrent 32/64/192 core Epyc line of OoO chips...How many strings does a Bass play with in water? =3
by Joel_Mckay
7/20/2026 at 4:16:47 AM
EPIC is Explicitly Parallel Instruction Computing; nothing to do with the AMD trademark word, EPYC.by shrubble
7/20/2026 at 4:40:38 AM
Indeed, the post does not conflate the two... look closer. =3by Joel_Mckay
7/19/2026 at 11:53:41 PM
Itanium is yet another example of how difficult it is to fight software ecosystem momentum.It had potential but until you have an easy transition path, few will consider it. Apple figured that out during the 68K > PPC > X86 > ARM transitions.
by ColdStream
7/20/2026 at 1:10:13 AM
Itanium 1 had hardware x86 compatibility and Itanium 2 had software x86 emulation just like Apple. It wasn't enough.by wmf
7/20/2026 at 2:29:12 AM
Right.PPC was significantly faster than m68k. Intel was significantly faster again. Finally ARM was yet again.
(ARM and Intel especially on laptops)
They could pay the emulation price and still come out equal or usually on top. Even when equal there were often other benefits, like reduced heat.
As I remember hearing Itanium was a dog with x86 code. It never got fast enough to compete let alone supplant it during emulation right?
So it wasn’t (meaningfully?) faster on recompiled code or new code. It wasn’t faster on new code. But it cost way more.
Not a winning combination.
by MBCook
7/20/2026 at 4:07:32 AM
IIRC Itanium was faster on natively compiled HPC code, presumably because it had more FPUs and larger cache than the Pentium 4. But in general it probably wasn't a good upgrade and Intel would have had to sandbag x86 a lot to force Itanium adoption.by wmf
7/20/2026 at 4:37:34 AM
The other problem was that the Pentium 4 itself was kind of a dog, and then the Pentium M and Core lines started making that look bad, by which point the Itanic passengers could feel their socks getting wet and a growing concern about there not being enough life boats.by AnthonyMouse
7/20/2026 at 5:37:58 AM
Only because AMD exists it wasn't enough. Had Intel been free to decide where x86 goes, the Itanium story would be much different.Apple succeeds on their hardware transitions exactly because there isn't a clones market.
by pjmlp
7/20/2026 at 12:13:20 AM
Agreed, everyone has a pet architecture that fascinates them, but if its so good only 5 people can code for the target... it is e-waste within a year.Apple is an exception as it has always had a walled-garden ecosystem with the OS, so can force shifts in architectures unlike most companies. The M3/M4 Pro series with unified GPUs is probably the best design on the consumer market right now, but people are not leveraging it as much as they would have in other ecosystems.
Have a wonderful day =3
by Joel_Mckay
7/20/2026 at 5:39:17 AM
Apple is the survivor of another era, people keep forgetting that x86 clones are the exception of how computer market used to be, not the other way around.by pjmlp
7/20/2026 at 2:31:04 AM
The walled garden is an advantage of some sort, but they’ve also never made a bad call on switching architectures.We don’t know what it would look like if Apple had chosen something that ended up like Itanium. It’s quite possible they don’t have the clout to pull it off.
(Maybe they should have gone Intel instead of PPC, but both were significantly better than m68k at that point)
by MBCook
7/20/2026 at 4:59:28 AM
> they’ve also never made a bad call on switching architectures.It seems more like they have a high willingness to switch architectures.
Predicting which one is better in the year that you do it or the couple of years after isn't that hard. The question is, where is it going to be in a decade or two?
Each time they picked a huge company you wouldn't have expected to fail in that year, first IBM, then Intel. But that's the problem with huge companies, after a few years on top they tend to get complacent and stagnate.
Now Apple itself is the big company, but it remains to be seen if they're not still going to end up on e.g. RISC-V within the next ten years or so.
by AnthonyMouse
7/20/2026 at 2:51:24 AM
>We don’t know what it would look like if Apple had chosen something that ended up like ItaniumApple made a few mistakes, but mostly by trying to compete with doomed hype markets like AR/VR.
Some also ponder what the ecosystem would look like today if the Windows NT kernel had stayed on RISC like initially planned.
The "What if __ ?" universe are fun to imagine, but ultimately less important than the "What now?" universe we live in. =3
by Joel_Mckay
7/20/2026 at 5:41:30 AM
I would already be happy if Windows NT had been more serious about UNIX compatibility.I only bought that Linux Unleashed book in 1995's Summer, because Windows NT wasn't good enough for doing university assignments at home, which used a mix of DG/UX and Solaris on the campus.
by pjmlp
7/20/2026 at 3:20:00 AM
Oh sure, Apple has TONS of mistakes if we look overall.I was only talking about CPU transitions.
by MBCook