Back
Vx puts heterogeneous memory and topology into the type system
SiTech AI Team3 min read

Vx puts heterogeneous memory and topology into the type system

Vx is a systems programming language for heterogeneous computing that puts CPU, GPU, NPU and accelerator memory into type checking, with explicit transfers and machine-aware placement.

Memory location is part of the type

Vx treats heterogeneous computing as a language-level concern rather than leaving accelerator placement to an opaque runtime. A tensor pinned to NPU high-bandwidth memory has a different type from the same tensor in host DRAM, and moving between those locations requires an explicit transfer. On Apple unified memory, that transfer compiles to almost nothing, but it remains visible in the source so data locality can be verified without profiling a binary.

The compiler checks for invalid host access to device pointers, pinned NPU values entering host expressions, and placements whose working set exceeds the target memory space. Other rules cover buffers read before an asynchronous transfer is visible, use-after-move errors, transfers without a declared path between memory spaces, and differentiation through regions without a defined adjoint. An SMT prover handles asynchronous transfer visibility contracts.

Machine files define hardware limits

Instead of hard-coding a cost model, Vx reads a machine file describing the memory hierarchy and interconnect of a particular part. The compiler uses that information to admit or reject placements. Unit conversions use exact integers: SI prefixes are decimal, while IEC prefixes are binary, preserving the meaning of values such as GB and GiB.

The repository includes machine files for H100, H200, B200, A100, MI300X, Apple M4 and multi-GPU nodes. The files cite their sources and identify figures that remain unverified.

A parallel frontend and multiple backends

Every symbol, nominal type and monomorphized variant in Vx's data-oriented frontend is represented as a flat 256-bit identifier. A nominal type system and mandatory boxing for recursive types decouple modules, allowing the compilation pipeline to run across cores without a query engine or lock contention. Compilation traverses flat arrays rather than pointer-chased trees.

The test suite asserts that serial and parallel builds emit byte-identical MLIR at benchmark scale under several thread configurations. This guarantee applies to frontend output; downstream processing belongs to LLVM. Backends cover x86-64 and arm64 CPUs, NVIDIA GPUs, Apple AMX and ANE, and distributed manifest-driven remote regions. Vendors can extend Vx through MLIR pass plugins rather than patching the compiler.

Built for fixed, performance-sensitive workloads

Vx is positioned for software that must be correct and fast across different kinds of silicon. The project acknowledges that its ahead-of-time, data-oriented and statically regioned design is less suitable for workflows that mutate architecture, inspect tensor shapes and continue dynamically while those decisions are still changing. Its installation options include macOS on Apple Silicon and Linux x86_64, while accompanying guides cover the language, topologies, memory placement and machine files.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.