7/23/2026 at 7:20:11 PM
Never really liked Geekbench. Synthetic benchmarks are easy to use and easy to understand but borderline useless. 1000 is better than 990 but it literally doesn't mean anything at all. Just an example: the 2017 Intel Xeon gets more points than an Apple M1 on multicore but everyone knows that they are incomparable. And no one should get a 2017 Intel Xeon because it's scoring higher yet that's the whole point of Geekbench, higher score is better.Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
Some good benchmark options here https://hwbot.org/benchmarks
https://www.pcgamingwiki.com/wiki/List_of_games_with_built-i...
by thrownawaysz
7/23/2026 at 7:56:15 PM
> the 2017 Intel Xeon gets more points than an Apple M1 on multicore but everyone knows that they are incomparableWhy are they incomparable? If I want to run a highly parallel task surely this tells me which one to use?
> Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
Pretty sure the internal tests it uses are similar to these
by saagarjha
7/23/2026 at 8:17:09 PM
In the past, they have provided information on which workloads they use.For instance, the file compression workload on version 6:
This workload compresses and decompresses the Ruby 3.1.2 source archive (a 75 MB archive with 9,841 files) using different compression codecs (such as LZ4 and ZSTD). It also verifies the files using the SHA1 (Secure Hash Algorithm 1) function.
by GeekyBear
7/23/2026 at 8:24:14 PM
It looks like they have the workloads for version 7 as well:CPU: https://www.geekbench.com/doc/geekbench7-cpu-workloads.pdf
GPU: https://www.geekbench.com/doc/geekbench7-gpu-workloads.pdf
by redsky17
7/23/2026 at 8:43:36 PM
Thanks for finding that.> The Clang workload uses the Clang compiler to compile the Lua interpreter, a popular open-source language interpreter.
People really shouldn't call Geekbench a synthetic benchmark anymore, since it has been using real world workloads for years now.
by GeekyBear
7/23/2026 at 9:08:28 PM
If memory serves, it has been since version 4 or 5 that they started including real workloads. I agree that it shouldn’t be considered synthetic any more.by redsky17
7/23/2026 at 11:17:20 PM
Benchmarks like 7Zip compression/decompression, LAME encoding, Blackmagic RAW video encoding, Cinebench, x265 encoding, Blender encoding, Y Cruncher, and any of the built-in video game benchmarks make much more sense than Geekbench.
I don't think so. First, Geekbench already shows sub scores for different types of applications.[0] Second, nearly no one uses CPU rendering for Cinebench and Blender. They're mostly GPU work. Video encoding/decoding work is mostly done by the media engine or GPU on Apple Silicon. Third, CPU benchmarks tend to correlate. If a CPU is faster in one thing, it's more likely to be faster in another. Therefore, a comprehensive score like what Geekbench and SPEC provide is valuable.Geekbench is extremely good at showing general CPU performance, especially ST. It's also highly correlated with SPEC at nearly 1:1 in terms of scores as shown by Nuvia before they were purchased by Qualcomm.[1]
[0]https://browser.geekbench.com/v6/cpu/18801240 scroll down
[1]https://medium.com/silicon-reimagined/performance-delivered-...
by aurareturn
7/24/2026 at 12:18:52 AM
While I definitely agree about video decoding, encoding still is a valuable measure.With inter-frame delivery codecs the hw encoder always focuses on speed, with a lean towards file size as well. It makes sense for most consumer use cases. 15x faster perhaps sure, but you might end up with a file almost 10x larger while still lower quality than a decent CPU encoding. Also the more off the beaten path of rec709 and rec2020 you go the less certain you can be that the hw encoder supports things.
by miladyincontrol
7/23/2026 at 9:26:49 PM
Benchmarks are like the SP500 index or IQ measurements: They distil multiple variables into one number. In doing so you lose details, but do get a useful measure.Yes, they are imperfect, but they do broadly measure how fast a processor is, and can be used for comparison.
by RachelF
7/24/2026 at 8:07:38 AM
>Never really liked Geekbench....As other have stated and perhaps worth pointing out clearly incase people don't get it. If you dislike Geekbench for those reasons, it is highly likely you don't know what Geekbench is testing in the first place.
Personally I think Geekbench is a pretty damn good consumer testing benchmarks. I just wish there is something similar for Server testing PHP, Ruby, JVM, MySQL and Postrges etc.
Can't wait to see all the M1 to M6 results. Along with AMD Zen 6.
by ksec
7/23/2026 at 10:02:13 PM
Borderline useless? Is it just that the numbers are unit-less?You can extract meaning from them! Find your current computer, find the target computer, calculate the percentage difference. 1000 is 1% better than 990, so it indeed doesn't mean anything at all.
by conradev
7/23/2026 at 8:43:04 PM
What you want exists, it's called SPECint/SPECfp.by BoingBoomTschak
7/23/2026 at 11:09:48 PM
Geekbench 5 and 6 both correlated with SPEC scores at nearly 1:1. SPEC is the industry standard for CPU benchmarks. So in a way, Geekbench does what OP suggests which is a general measurement of performance. If you want to see specific areas, Geekbench shows sub scores.Here's research done by Nuvia before they were bought by Qualcomm: https://medium.com/silicon-reimagined/performance-delivered-...
by aurareturn
7/24/2026 at 8:32:09 AM
What compiler is used by GB? Why can't I change it? Imagine an ICC-like scandal. Trusting binary-only benchmarks to compare hardware simply is foolish.by BoingBoomTschak
7/24/2026 at 4:36:30 PM
I don't think you've put forth a coherent argument here.SPEC CPU is distributed in source form, to be compiled by the end user. There are restrictions on what kind of compiler options and other games you can play when making an official submission to SPEC, which Intel (and others) get around by publishing "estimated" SPEC scores and using those for marketing claims. In practice, the amount of bullshit in Intel's claimed SPEC scores is constrained only by what Intel's legal department is willing to risk getting sued over.
Geekbench is distributed in binary-only form with all binaries for all platforms originating from the author of Geekbench. All the major chip designers pay to see the source code and participate in the development, but the final decisions about what code goes in and how it's compiled are made by the author of Geekbench, not by the chip vendors. They just get to lobby, and AMD has every opportunity to call out Intel and Geekbench should the compiler options for the x86 builds be unfair.
Intel has a long history of cheating on SPEC, but cheating on Geekbench is more difficult because of the necessity of runtime binary patching and they didn't publicly try until this year, with limited success and lots of negative press: https://www.geekbench.com/blog/2026/03/geekbench-6-and-intel...
by wtallis
7/24/2026 at 3:30:59 PM
Yes, for single-threaded benchmarks Geekbench 5 and 6 are very well correlated with the SPEC scores and with other single-threaded CPU benchmarks.On the other hand, whether the Geebench multi-threaded benchmarks are meaningful is questionable. Especially Geekbench 6 has a very poor scaling to many threads, which shows results that are quite uncorrelated with many practical workloads.
Moreover, the multi-threaded benchmarks are too short to reach a steady-state, so they show extremely optimistic throughputs for CPUs with poor cooling, like those from smartphones and laptops. Thus with Geekbench many smartphones and laptops appear to be competitive in multi-threaded performance with desktops or servers, while in real workloads their performance would be pathetic (because as soon as they would overheat, their performance would drop drastically).
Like any benchmarks, Geekbench can be gamed by CPU vendors.
For example, because already the older versions of Geekbench included secure hash computations, the Arm Aarch64 CPUs obtained good scores, because since 2012 they had special instructions and hardware for computing SHA-1 and SHA-256.
The x86-64 CPUs did not have such instructions, so they had bad scores. This prompted Intel to add a SHA ISA extension, but they added it only to the Atom CPUs, because only those competed with Arm CPUs and people compared their Geekbench scores.
For many years, among the Intel CPUs only the small and cheap Atom CPUs had SHA instructions, while their big and expensive desktop/laptop/server CPUs did not have SHA instructions, because at that time nobody compared their Geekbench scores.
This has changed only after the launch of AMD Zen, which included the SHA extension. Then, eventually, after the years during which Intel failed to progress beyond the Skylake derivatives, Intel finally added the SHA extension to all its CPU series, starting with Ice Lake, and since then all the x86-64 CPUs have it.
by adrian_b
7/23/2026 at 8:48:55 PM
Since the loss of Anandtech, have any of the other tech websites started publishing SPEC results for systems they test?by GeekyBear
7/23/2026 at 11:14:40 PM
David Huang:by aurareturn
7/23/2026 at 11:43:37 PM
What a great resource!Thanks
by GeekyBear
7/23/2026 at 9:51:52 PM
https://chipsandcheese.com/by BoingBoomTschak
7/24/2026 at 8:41:04 AM
By the way, here's an interesting and very related comparison by them: https://chipsandcheese.com/p/evaluating-geekbench-6by BoingBoomTschak
7/23/2026 at 9:45:39 PM
Phoronix, I've at least seen it on Epyc Reviews iirc.by ysleepy
7/24/2026 at 8:01:27 PM
To be honest I always got the sense that it had a sort of pro-Apple bias that doesn't seem to be reflected in real world applications.by captainbland
7/23/2026 at 8:47:56 PM
AV1, Opus, Whisper, Jolt Physics... did you read the article? Are these not real?by wmf
7/23/2026 at 7:55:48 PM
a 2017 Xeon can absolutely smoke an M1.. (double the FP64 units, AVX-512; huge per core caches, boatloads of I/O lanes and RAM...).. but it chugs power like it's not even funny.
I do like synthetic benchmarks as one input signal when evaluating a CPU; otherwise we go back to stupid numbers like frequency... which also never made sense because they had 1.8GHz AMD Athlon CPUs outperforming 2.4GHz Pentium 4's in 2002-3...
Doing a lot of real-world application testing is good, but there's always some you miss, and the worst part about it is that people start to lose patience when assessing.. How is a cinebench comparing to doing gamedev in UE5?
How is doing gamedev in UE5 comparing to doing gamedev in Snowdrop?
How does doing gamedev in Snowdrop compare to doing virtualisation (different CPUs handling that particular task better than others).
there's so many dimensions that it will always be true that there's not enough testing.
I'm not saying we shouldn't have rigorous testing like you say, in fact, what I'm actually saying is that "single number→would you like to go deeper" is a better pipeline than an excel spreadsheet that goes on for 10 pages (which still skips a bunch of nuance) and still better than one that goes on for 500 pages with a broad spectrum of topics.
The reason I'm so bitter about this is because outlets like LTT and Gamers Nexus seem to choose their games randomly (based on popularity I guess?) and so all the titles I worked on kinda got smeared by my "company"- but our games used hugely different game engines that work differently (Dunia vs Snowdrop handle CPUs... VERY differently. Snowdrop is built for multi-core. Anvil works best on a single core, and Dunia has trouble going passed 4 cores- so on large systems with multiple NUMA zones the performance is.. worse.
So the "battery of tests" always gave us a shitty score and reviewers just moved on believing themselves to be comprehensive arbiters of truth regarding performance and giving verdicts regarding the performance of games as a broad topic (which people take and don’t dig deeper themselves) and without regard to the utility of what they’re saying (500+ fps on CSGO isn't even renderable for example).
People might have been better served by looking broadly at how powerful the CPUs actually are and then digging in for their use case once they whittle a few down, because then also: developers will actually make software that tries to use the features on offer instead of assuming people will just buy CPUs that work better on their workloads.
It's a kind of "to big to fail" mentality where incumbents start being able to direct CPU sales based on those CPUs being optimised for their workloads. It's self-reinforcing.
by dijit