News6 hours ago

Go's Portable SIMD Beat Plain Go by 10x on Our M4 and Lost to It When Emulated

Go 1.27's experimental simd package ran a float32 dot product 10.6x faster than a plain loop on our M4 Pro. Forced into emulation, the same code ran 1.9x slower than the loop it replaced.

The WJS Desk

Sep 28, 2026 · 6 min read

Photo by Jimmy Chan on Pexels

On September 24 the Go team published a long post explaining the new simd package, the portable SIMD layer that shipped behind GOEXPERIMENT=simd in Go 1.27. It hit 407 points on Hacker News. It also contained zero benchmark numbers, which for a feature whose entire reason to exist is speed felt like a gap worth filling.

So we downloaded Go 1.27.1, wrote a float32 dot product four ways, and ran each one on an Apple M4 Pro. The short version: written carefully, portable SIMD ran 10.6x faster than a plain Go loop. Written the way the blog post's own example is written, it ran 3.7x faster. And with hardware SIMD switched off, the "emulation" path ran 1.9x slower than the plain loop it was supposed to replace.

What actually shipped

Go 1.27 (August 2026) has two experimental packages, both enabled only when you build with GOEXPERIMENT=simd:

  • simd/archsimd, the architecture-specific layer that started on amd64 in Go 1.26. The 1.27 release notes say it now also covers arm64 Neon and WebAssembly, at 128 bits, with 256 and 512 bit types on some amd64 chips. Non-portable by design.
  • simd, the new portable layer. Its vector types (Float32s, Int8s and so on) have no fixed width. The width is chosen once at program start: 128, 256 or 512 bits depending on the CPU, or pure-Go emulation where no hardware support exists.

Per the post by David Chase and Junyang Shao, the compiler makes this cheap by cloning every function that mentions a simd type into size-specialised copies and hoisting the dispatch to callers. Arm64 SVE is planned for Go 1.28. There is no reduction op yet: ReduceSum "will appear in the next release", so today you sum the lanes yourself.

Without the experiment flag nothing works at all. We confirmed that a module importing simd fails with build constraints exclude all Go files, so this cannot leak into a normal build by accident.

How we tested it

Setup: Go 1.27.1 darwin/arm64, Apple M4 Pro, macOS 26.5.1, go test -bench with -count 5, summarised with benchstat. On this chip the package reports 4 float32 lanes (128-bit Neon) and simd.Emulated() returns false.

Four implementations of the same dot product:

  • Scalar: the obvious for i := range x { s += x[i] * y[i] }
  • Scalar4: the same loop with four independent accumulators, which is what a performance-minded Go programmer writes today
  • SIMD: the blog post's innerProduct example, one vector accumulator and MulAdd
  • SIMD4: the same thing with four vector accumulators

The SIMD4 inner loop, which is the only version we would ship:

var a0, a1, a2, a3 simd.Float32s
n := a0.Len()
for ; i < len(x)-4*n+1; i += 4 * n {
    a0 = simd.LoadFloat32s(x[i:i+n]).MulAdd(simd.LoadFloat32s(y[i:i+n]), a0)
    a1 = simd.LoadFloat32s(x[i+n:i+2*n]).MulAdd(simd.LoadFloat32s(y[i+n:i+2*n]), a1)
    a2 = simd.LoadFloat32s(x[i+2*n:i+3*n]).MulAdd(simd.LoadFloat32s(y[i+2*n:i+3*n]), a2)
    a3 = simd.LoadFloat32s(x[i+3*n:i+4*n]).MulAdd(simd.LoadFloat32s(y[i+3*n:i+4*n]), a3)
}

We checked all four against the scalar result on a 1,003-element input (deliberately not a multiple of the lane count) before timing anything.

What we measured

Median time per dot product, lower is better:

Implementation1,024 floats65,536 floats1,048,576 floatsvs Scalar at 1M
Scalar739.7 ns51.53 µs854.1 µs1.0x
Scalar4357.1 ns24.28 µs387.5 µs2.2x
SIMD (blog example)214.9 ns14.09 µs229.5 µs3.7x
SIMD481.08 ns4.974 µs80.51 µs10.6x
SIMD4, GODEBUG=simd=01.668 µs101.1 µs1.676 ms0.51x

Three things stand out.

The blog's example leaves most of the speed on the table. One accumulator means every MulAdd waits for the previous one to finish, so the loop is bound by FMA latency rather than throughput. Splitting it four ways made it 2.85x faster again. Measured against the four-accumulator scalar loop a careful Go programmer already writes, the example's win shrinks from 3.7x to 1.7x.

Emulation is slower than plain Go. With GODEBUG=simd=0, which forces the same emulated path a CPU with no supported SIMD would get, the SIMD4 version took 1.676 ms at 1M floats. That is 1.9x slower than the naive loop and 4.3x slower than Scalar4. Unrolling made no difference in emulation, both SIMD versions landed at about 4.7 GiB/s.

The results are not bit-identical. On our 1,003-element check the scalar loop returned 5.574726 and the single-accumulator SIMD version returned 5.5747237. That is normal floating-point reassociation, but any golden test that compares float32 output exactly will fail the day you switch.

Gotcha: GODEBUG=simd=512 does not mean "use 512 bits if you can". On our 128-bit Neon chip it panicked in simd.init before main ran, and so did the +512 form and 256. The panic message also calls the setting gosimd, while the blog and the source use simd. Do not put a width in a shared deployment config unless every machine it reaches supports it.

What we did not test: real x86 hardware with AVX2 or AVX-512, which is where the 256 and 512 bit paths live and where the post implies the biggest wins are. Under Rosetta on the same Mac, the amd64 build reported 4 lanes, so that is not a stand-in for real x86 numbers. We also cross-compiled for wasip1/wasm and linux/riscv64 and both built cleanly, but we did not run them. Our largest input is 8 MB of data, which fits in this chip's cache, so these are compute numbers, not memory bandwidth numbers.

What the threads said

The Hacker News thread was mostly warm, and the few people with numbers roughly agree with ours. ImJasonH built a browser demo that swaps colours in an image under wasm, and reports portable SIMD about 11% slower than archsimd, with both around 5x faster than no SIMD. karolist reports about a 30% speedup on foreground estimation for image cutouts, while admitting the algorithm is not tuned yet.

The sharpest dissent came from melodyogonna, who argued that vector-length-agnostic code can "end up with code that performs much worse than the scalar alternative on some platforms", and that specialising per platform with portable code as the fallback is usually better. Our emulation row is that scenario, measured. cryptolobster asked how much emulation ends up in hot paths before SVE lands. On a CPU with no supported SIMD, the answer from our numbers is: all of it, at half the speed of the plain loop.

On the other side, mshockwave pointed out this is the first portable SIMD design they have seen that makes variable-length hardware like SVE and RISC-V's vector extension straightforward to target, and janwas from Google's Highway library said they had given the Go team advice on exactly that choice. physicsguy raised the other gap: most people want autovectorisation, not an API. typical182 linked an in-flight compiler change stack (CL 791740) doing exactly that, though nothing there is promised for a release.

Should you use it

If you run Go on arm64 or modern amd64 and have a hot numeric loop that you currently push through cgo or hand-written Go assembly, yes, start experimenting now. A 10x gain from pure Go with no assembly file is real, and the portable API meant our benchmark ran unchanged under Rosetta and cross-compiled for wasm without a single build tag.

But treat three things as requirements, not suggestions. Write multiple accumulators, because the idiomatic one-accumulator loop gets you a third of the speed. Benchmark with GODEBUG=simd=0 if any of your targets might lack hardware support, because the fallback loses to the loop you already have. And keep the scalar version around behind a build tag, since the whole package sits behind an experiment flag and the release notes say the related archsimd API is not yet stable.

Portable SIMD is only portable in the sense that it compiles everywhere. It is fast in far fewer places.

Your turn

If you have tried GOEXPERIMENT=simd on real x86 hardware, we would love the number we could not get: what does simd.Float32s{}.Len() report on your AVX2 or AVX-512 box, and how does that SIMD4 loop compare to the scalar one? Paste the benchstat line.

And if vendor performance claims with missing numbers are your thing, we did the same exercise when NVIDIA announced Rust GPU kernels without benchmarks and went digging through the repo for them instead.

Share

Go's new portable SIMD made our dot product 10.6x faster on an M4. Switch the hardware path off and the same code ran 1.9x slower than a plain for loop. #golang #SIMD #Performance

Never miss a ship

The best stuff that shipped this week, delivered every Thursday. Free, no spam. We read all the boring stuff so you get the fun parts.

Keep reading