Go archsimd Tutorial: Use SIMD in Go 1.27
Learn how Go 1.27's experimental archsimd package works, enable GOEXPERIMENT=simd, write vector code, handle CPU features, and benchmark it safely.
On this page
Go 1.27 gives developers a new way to reach CPU vector instructions without dropping into Go assembly: the experimental simd/archsimd package. The October 2, 2026 Go team article explains how the API now reaches amd64, arm64, and WebAssembly targets, while the portable simd package provides a higher-level alternative. This Go archsimd tutorial builds a small vector-add program, enables the experiment correctly, handles CPU feature checks, and shows where architecture-specific SIMD actually makes sense.
Why Go 1.27 makes SIMD easier to try
Single Instruction, Multiple Data (SIMD) lets one CPU instruction operate on several values stored in a vector register at once. A normal loop might add one pair of numbers per iteration, while a SIMD operation can add several pairs together in parallel when the data and instruction set allow it. Go previously required assembly to make direct use of this kind of hardware acceleration, but Go 1.26 introduced the experimental simd/archsimd API for AMD64 and Go 1.27 extends the experiment to ARM64 and WebAssembly.
There are two related APIs in Go 1.27. The simd package is the portable layer: it hides vector width and uses hardware instructions where available, with emulation where necessary. The simd/archsimd package is lower-level and exposes architecture-specific vector types and operations, so it is the one to use when a particular CPU instruction or layout matters. Both remain experimental and require GOEXPERIMENT=simd when building.
Check Go 1.27 before enabling the experiment
Start by checking the compiler version and architecture you are actually using. The simplest commands are:
go version go env GOOS GOARCHYou want Go 1.27 or a later release and a target architecture supported by the experiment. Go's 1.27 documentation lists AMD64, ARM64, and WebAssembly support for the experimental architecture-specific SIMD work, while the exact instruction set available still depends on the processor running the program.
Do not treat GOEXPERIMENT=simd as a permanent project-wide requirement unless the whole application is deliberately built around experimental SIMD support. It changes the build configuration and exposes APIs that are not covered by Go's normal compatibility promise. For a first experiment, enable it only for the command or development environment where you are testing the code.
Enable GOEXPERIMENT=simd before importing archsimd
The simd/archsimd package is excluded unless the SIMD experiment is enabled. On Linux or macOS, set the environment variable for the command like this:
GOEXPERIMENT=simd go run.On Windows PowerShell, the equivalent is:
$env:GOEXPERIMENT="simd" go run.If you run go run. without the experiment enabled, code importing simd/archsimd will not behave like an ordinary stable-library import because the package is guarded by the experimental build configuration. The Go source itself confirms that the SIMD package is built under the goexperiment.simd constraint.
Build a small vector addition example
Now create a small program that adds four float32 values at a time. This example deliberately targets AMD64 because the architecture-specific API is designed around the instruction capabilities of the target CPU, and it makes the CPU-feature requirement easy to demonstrate.
//go:build goexperiment.simd && amd64
package main
import (
"fmt"
"simd/archsimd"
)
func main() {
if!archsimd.X86.AVX() {
fmt.Println("AVX is not available on this CPU")
return
}
a:= []float32{1, 2, 3, 4}
b:= []float32{10, 20, 30, 40}
out:= make([]float32, 4)
va:= archsimd.LoadFloat32x4(a)
vb:= archsimd.LoadFloat32x4(b)
va.Add(vb).Store(out)
fmt.Println(out)
}The important part is not the size of the example but the data flow. LoadFloat32x4 loads four floating-point values into one SIMD vector, Add adds the corresponding lanes, and Store writes the resulting vector back into a normal Go slice. The API documentation identifies Float32x4 as a 128-bit vector containing four float32 values, while the addition operation maps to the CPU's vector addition instruction on AMD64.
Go Packages
Run the program and check the result
Save the program as main.go and run it with the experiment enabled:
GOEXPERIMENT=simd go run.The expected output is:
[11 22 33 44]That output gives you a useful first check: four scalar additions have been expressed as one vector operation in the Go source. It does not, however, prove that the SIMD version is faster for your real application. Memory access, loop overhead, the size of the data set, compiler optimizations, and the CPU's available instruction set can all determine whether vectorization produces a measurable improvement.
Turn the four-value example into a real loop
A useful SIMD function normally processes a large slice rather than exactly four values. The basic structure is a vector loop followed by a scalar tail for elements that do not fill a complete vector. Go 1.27 adds LoadFloat32x4Part and StorePart specifically for cases where the remaining slice is shorter than a full vector, but a conventional vector loop plus scalar remainder is often easier to understand when starting out.
//go:build goexperiment.simd && amd64
package main
import "simd/archsimd"
func addFloat32(a, b, out []float32) {
if len(a)!= len(b) || len(out) < len(a) {
panic("slice lengths do not match")
}
if!archsimd.X86.AVX() {
for i:= range a {
out[i] = a[i] + b[i]
}
return
}
i:= 0
for; i+4 <= len(a); i += 4 {
va:= archsimd.LoadFloat32x4(a[i:])
vb:= archsimd.LoadFloat32x4(b[i:])
va.Add(vb).Store(out[i:])
}
for; i < len(a); i++ {
out[i] = a[i] + b[i]
}
}The scalar tail matters because a slice does not have to contain a multiple of four values. A seven-element input can process four values with SIMD and then finish the remaining three normally. This is also why a SIMD implementation should not blindly load a full vector at every index: the documented full-width load expects enough elements to exist in the slice. Go Packages
Check CPU features before using optional instructions
The feature check is more than defensive programming. The Go team warns that executing an instruction on hardware that does not support it can terminate the process with an illegal-instruction signal. The archsimd package therefore exposes runtime checks such as archsimd.X86.AVX(), AVX2(), and AVX512(), allowing the program to choose an appropriate implementation for the machine it is actually running on.
This becomes especially important when you move beyond simple addition. Different methods require different CPU features. The package documentation identifies the instruction-set requirement for individual operations, so you should check the requirement for the operation you are using rather than assuming that every SIMD operation available in the API works on every AMD64 processor.
Do not skip the feature check: if an instruction requires a CPU feature that the machine does not provide, the resulting binary can fail at runtime rather than merely becoming slower.
Use the new SIMD API for more than arithmetic
Vector addition is deliberately simple, but archsimd exposes arithmetic, comparisons, bitwise operations, conversions, permutations, masks, loads, and stores. Go 1.27 also replaces the older type-specific reinterpretation approach with composable conversions such as ToBits() and BitsToFloat32(). That matters for byte-processing and numeric algorithms where the same bits need to be viewed through different element types without performing a real data conversion.
Masks are another important part of the API. A comparison such as x.Greater(y) produces a mask, which can then be used with operations such as Masked or IfElse. The compiler can recognize some combinations and lower them into more efficient masked instructions when the required CPU features are known.
Know when archsimd is the wrong layer
Architecture-specific SIMD is useful when you need direct control over vector operations, but it is not automatically the right choice for every performance-sensitive loop. The simd package introduced in Go 1.27 deliberately hides fixed vector widths and exposes a portable set of operations that can use hardware SIMD where supported or emulate the operation where it is not. That makes it a better fit when the same algorithm needs to run across different CPU architectures without maintaining architecture-specific code paths.
There is a practical trade-off. archsimd can expose instructions that do not fit the portable API, but those operations tie your implementation more closely to the target architecture. The Go team specifically describes archsimd as lower-level infrastructure and says its available types and operations depend on the target architecture. It also advises against exposing these SIMD types in public APIs because they are not portable.
Benchmark the complete workload instead of the vector instruction
A SIMD loop can look dramatically different from scalar Go without producing a meaningful application-level speedup. If your workload spends most of its time loading data from memory, calling other functions, or handling small slices, the vector arithmetic may not be the bottleneck. A useful benchmark should therefore compare complete functions using realistic input sizes rather than timing only the Add operation.
Go's normal benchmark framework is enough for this experiment. Create one benchmark for the scalar implementation and another for the SIMD implementation, keep their input data equivalent, and run them repeatedly with the same build configuration. Then test more than one input size because a SIMD path that helps a large batch can add unnecessary setup or dispatch cost to a tiny operation.
GOEXPERIMENT=simd go test -bench=. -benchmem./...The result you want is not simply a smaller number in a benchmark table. Look for whether the improvement survives realistic input sizes and whether memory allocation, tail handling, or feature dispatch becomes the new bottleneck. The Go team's own SIMD examples emphasize CPU checks, bounds-check behavior, and keeping vectors in registers because those details can determine whether an apparently efficient SIMD implementation actually remains efficient after compilation.
Test the same experiment on WebAssembly or ARM64
Go 1.27 extends the architecture-specific experiment beyond AMD64. The Go team lists ARM64 with Neon support and WebAssembly with 128-bit SIMD, while AMD64 remains the broadest implementation with AVX, AVX2, and several AVX-512 extensions. That makes the new API relevant to Apple Silicon and ARM servers as well as conventional Intel and AMD systems, although the exact operations available still vary by architecture.
The Go project provides direct test commands for these targets. For WebAssembly, its documented example uses:
GOOS=wasip1 GOARCH=wasm GOEXPERIMENT=simd go test simd/archsimd/...On an Apple Silicon machine, the Go team also documents testing an AMD64 build through Rosetta by setting GOARCH=amd64 together with the SIMD experiment. These commands are useful because they test the actual architecture-specific implementation rather than assuming that code written for one CPU family will behave identically everywhere.
Keep archsimd behind a small implementation boundary
A good production design is to keep the SIMD-specific code in a small internal package or implementation function and expose ordinary Go slices or application types to the rest of the program. That keeps architecture-specific vector types out of public APIs and gives you a natural place to maintain the scalar fallback. The approach also makes it easier to disable the experimental path when a compiler or target changes.
For a first project, vector addition is enough to verify the mechanics: enable GOEXPERIMENT=simd, load vectors, perform the operation, store the result, check the CPU feature, handle the tail, and benchmark the complete workload. From there, move to a real hot loop such as numeric filtering, image or byte processing, or another data-parallel operation where several independent values receive the same computation. Go 1.27 has made the low-level path much more approachable, but the experiment is still experimental, so the next step is to measure whether your workload actually benefits before building an application around it.
Written by


