
Test patch removes byte shuffling from Go netip IPv4-to-IPv6 mapping
A test compiler rewrite turns Go's netip.AddrFrom16(ip.As16()) mapping pattern into a direct internal field copy, producing the same short assembly as the 0.88 ns standard library solution.
Why the mapping pattern is slow
Go's netip.Addr provides an Unmap() method that returns the unwrapped IPv4 address from an IPv4-mapped IPv6 address. There is no Map() or To6() method for the reverse operation. Go maintainers have rejected adding one, directing users to netip.AddrFrom16(ip.As16()) and hoping the compiler will optimize it.
Internally, netip.Addr stores an IP address as a 128-bit value and uses an additional field to encode its family and zone. A standard library implementation of To6() could change that field for IPv4 addresses. An external helper must instead convert the address to a 16-byte array and back.
Benchmark results
In benchmarks using Go 1.27.1 on Linux and AMD's Ryzen 5 5600X 6-Core Processor, the standard library solution took 0.8775 ns per operation. The safe helper took 7.137 ns, while an unsafe implementation that accessed the same internal state through a proxy took 0.8682 ns. The helper was therefore about eight times slower.
The standard library and unsafe versions produced nearly identical assembly. Both checked the address family and replaced the internal family value when processing IPv4. The safe helper generated more instructions because it packed the address into a 16-byte array, copied the array and unpacked it again. As of Go 1.26.8, the compiler did not remove this sequence.
A compiler-level experiment
The experiment rewrites the exact netip.AddrFrom16(ip.As16()) pattern during the compiler's noding phase. It replaces the calls with a netip.Addr structure that copies the original address and assigns the IPv6 family value. The rewrite must happen during noding because the earlier type-checking phase prevents access to the structure's unexported fields.
A custom build with the rewrite generated the same short assembly as the direct implementation, and tests passed for net/netip and the helper package. The analysis says Go maintainers are unlikely to accept this approach because it depends on the internal layout of netip.Addr and gives the noder responsibility beyond faithfully translating the type-checked syntax tree.
The proposed SSA optimization
The more general approach targets the compiler's generic SSA phase, specifically the memcombine pass. The proposed rewrite rules redirect loads through memory moves, forward stored values to later loads and cancel pairs of byte-swap operations.
For each 64-bit half of the address, these transformations remove the temporary array, its copy and the unnecessary memory access. The higher half can also be forwarded across a store to a non-overlapping adjacent address. After simplification, each output half becomes a copy of the corresponding input half, leaving the address data unchanged while setting the required family value.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.