
Writing efficient C++ code: why data layout decides performance
Programmer Adam Sawicki explains why choosing C++ is not enough for a fast program: the result is decided by how data lies in memory, how the cache is used and how much allocation costs. The article first appeared in Polish in Programista in 2013.
Choosing C++ does not by itself make a program fast. In “Writing Efficient C++ Code”, programmer Adam Sawicki explains what decides whether code really uses the hardware: how data lies in memory and how often the processor waits for the cache. The text first appeared in Polish in Programista in 2013.
C++ alone does not buy speed
C++ is high-level enough for object-oriented design and ready-made containers, yet low-level enough that no virtual machine or garbage collector stands between the code and the operating system: memory is freed by the programmer.
Data-Oriented Design: data before algorithms
Data-Oriented Design, popular among game programmers, starts with the data: its layout in memory and the structures that hold it, with algorithms chosen afterwards. It does not ban classes, but warns against many small objects scattered in memory and linked by pointers: such layouts cause frequent cache misses and are hard to parallelise. A regular array, by contrast, is processed element by element and can be split across threads.
The memory hierarchy sets the price
On an example 3 GHz machine one arithmetic operation takes about 0.33 nanoseconds, a single cycle. A value in the L1 cache costs roughly 1 nanosecond, in the L2 about 4.7, and in RAM around 83 nanoseconds, the equivalent of 250 cycles. Data travels in 64-byte cache lines, so values used together should lie next to each other.

Contiguous structures such as a plain array or std::vector therefore beat linked lists, trees and graphs. The article sketches a “performance pyramid”: arithmetic is fastest, transcendental operations such as sine or division are much slower, a RAM access after a cache miss costs hundreds of cycles, dynamic allocation is expensive, and input and output is slowest of all.
AOS, SOA and compiler limits
A particle system shows the trade-off. An Array of Structures keeps position, velocity and colour together for each particle; a Structure of Arrays keeps a separate array per property, so the values used in one loop stay close and unused fields never enter the cache.
The compiler cannot fix everything. When a pointer may alias array elements, it must reload the value on each iteration, so copying it to a local variable or using __restrict helps. Moving a std::string out of a loop and clearing it shortened a test from 0.36 to 0.27 seconds, 25%, because the string keeps its buffer. Sawicki also notes that what a compiler cannot prove, the programmer can often arrange by hand.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.