⚡ Zig Guide LiveUnofficial
✓ Zig 0.17.0-dev.1454+5faa79730On an older Zig?

Floats

const std = @import("std");
const expect = std.testing.expect;

test "float widths" {
    const a: f16 = 1.5;
    const b: f32 = 1.5;
    const c: f64 = 1.5;
    try expect(a == 1.5 and b == 1.5 and c == 1.5);
    try expect(@sizeOf(f16) == 2 and @sizeOf(f64) == 8);
}

test "int and float do not mix implicitly" {
    const n: i32 = 3;
    const f: f32 = @floatFromInt(n);
    try expect(f == 3.0);

    // Truncates toward zero; checked in safety builds.
    const back: i32 = @intFromFloat(3.9);
    try expect(back == 3);
}

test "comptime_float is f128, not magic" {
    // Unlike comptime_int, which really is arbitrary precision,
    // comptime_float is backed by f128. Doing the arithmetic at compile
    // time buys you more bits, not exactness:
    const x = 0.1 + 0.2;
    try expect(x != 0.3);

    // The same sum in f64 is famously not 0.3 either.
    const y: f64 = 0.1;
    const z: f64 = 0.2;
    try expect(y + z != 0.3);

    // Compare with a tolerance instead of ==.
    try expect(std.math.approxEqAbs(f64, y + z, 0.3, 1e-12));
}

test "special values" {
    const inf = std.math.inf(f32);
    const nan = std.math.nan(f32);
    try expect(std.math.isInf(inf));
    try expect(std.math.isNan(nan));
    try expect(nan != nan); // NaN is not equal to itself
}

Zig has f16, f32, f64, f80, and f128. All are IEEE-754 (except f80, the x87 extended format).

No implicit int/float mixing

Neither direction happens on its own:

const f: f32 = @floatFromInt(n);
const n: i32 = @intFromFloat(f);   // truncates toward zero

This is deliberate. Implicit int-to-float conversion silently loses precision for large integers, and float-to-int silently truncates; both are common enough sources of bugs to be worth two extra words.

comptime_float is f128, not magic

This one catches people, and it caught an earlier draft of this page. Unlike comptime_int, which really is arbitrary precision, comptime_float is backed by f128. Doing arithmetic at compile time gives you more bits, not exact decimal:

const x = 0.1 + 0.2;
// x != 0.3, even at comptime

Floating point is still floating point. Compare with a tolerance:

std.math.approxEqAbs(f64, a, b, 1e-12)

NaN and infinity

std.math.nan(f32) and std.math.inf(f32) construct them; isNan and isInf test for them. Remember nan != nan. That is exactly why approxEqAbs exists and why == on floats deserves suspicion.