Repository navigation
allow integer types to be any range #3806
Description
Activity
- addedproposalThis issue suggests language modifications. If it also has the "accepted" label then it is planned.This issue suggests language modifications. If it also has the "accepted" label then it is planned.
on Nov 29, 2019 I'm worried this might result in a code explosion as different types are passed to a generic.
Idea: input to a generic could be "marked" for which properties are used.
e.g. if I have a pure-storage generic, likeArrayList, then the range of the members doesn't really matter for the implementation of that generic: the only property it cares about at runtime is the size of the type, the rest is only useful at comptime.Reacted by pixelherodev, Felix Queißner, zfl9, markus, EtienneParmentier and pancelor@andrewrk i cannot understand how did you get required size of
MyFancyTypestruct 'less that one byte ... leaving room for 235 more tags'.Since struct is a product type, MyFancyType has a total of
4 * 2 * 6 * 9 =432 possible states. Requiring at least⌈log₂(432)⌉ =9 bits to store them.Reacted by jibal, Borhan Hafez, LearnLoop, pancelor and Natalie Sepruko@Rocknest my math is wrong, thanks for the correction 👍
Reacted by Björn Roberg, Ethan Thoma, erdfern, LearnLoop, pancelor and Natalie Seprukowouldn't it be
@Type(TypeInfo.Int{...})whereTypeInfo.Intlooks like:pub const Int = struct { is_signed: bool, bits: comptime_int, from: comptime_int, to: comptime_int, };
because of the
@Typebuiltin?Yes except only
fromandtoare needed - the rest is redundant information. I'd call itstartandendto hint that the upper bound is exclusive. Good point, I forgot that@Intis deprecated.@Int(0, 1) == @Type(.{ .Int = .{ .start = 0, .end = 1}})@Typeis a bit more verbose though, for such a fundamental type as an integer. I'm not 100% sold on@Typeyet.should this be marked a breaking proposal as well?
Reacted by Andrew Kelley and pixelherodev- addedbreakingImplementing this issue could cause existing code to no longer compile or have different behavior.Implementing this issue could cause existing code to no longer compile or have different behavior.
on Dec 1, 2019 Yes except only
fromandtoare needed - the rest is redundant information.You still need the
bits. e.g. I might have an API that takes ausizebut only allows values between 1 and 40961is >= 0, so it's unsigned- log2(4096) = 12, so it's 12 bits
You're thinking of a different, already planned issue, which is zig keeping track of additional information beyond the type, for expressions and
constvalues, which will allow, for example, this test to pass:test "runtime value hints" { var a: u64 = 99999; // `a` is runtime known because it is `var`. const b: u64 = a % 256; const c: u8 = b; // allowed without @intCast because `b` has a runtime value hint }
i don't think so. take for example the WebSocket protocol, the opcodes are always 4 bits, but 0xB to 0xF are reserved for future use, so effectively they are not allowed.
i don't think so. take for example the WebSocket protocol, the opcodes are always 4 bits, but 0xB to 0xF are reserved for future use, so effectively they are not allowed.
That sounds like a usecase for a non-exhaustive enum:
enum(u4) { Continuation = 0, Text = 1, Binary = 2, Close = 8, Ping = 9, Pong = 10, _, }
This is very nice and fancy, but I would like to note that it might have a negative effect on performance. This is definitely more efficient memory wise, so for storing it (on disk) this is perfect. However for passing it around in code in this format it would mean that every time there is a calculation bit fiddling comes into play to extract the value and put it back.
Just something to think about.
However for passing it around in code in this format it would mean that every time there is a calculation bit fiddling comes into play to extract the value and put it back.
This is not correct, except with
align(0)fields (see #3802). Zig already has arbitrary sized integers. If you usei3, for example,@sizeOf(i3) == 1because@alignOf(i3) == 1.Also in many cases if you were to use
align(0)fields, the smaller memory use on the cache lines would cause more performance gains than lost to the extra bit shifting.Reacted by Chris Heyes and Dan Kubb42 remaining items
It seems worth observing that this offers a possible solution to the felt need in #17039/#14704. Andrew made a comment to that effect in the first of those issues, I wanted to add some notes about it here.
If and when integers are defined as a range, rather than a width, it could be worthwhile to make this legal:
for (u8) |byte| { // ... }
If I saw this I would have no question about what values
bytewill take, or the type, or the order in which they're iterated.If that's too weird, it could take a builtin:
for (@range(u8)) |byte| { // ... }
But I think that in a world where integers are ranges, it would be eloquent for them to coerce to an iteration of that range in a
forloop position. Anything involving a different order or stride would need to use awhileloop, this approach wards off the temptation to add 'just a bit more syntax', it treats an integer type as implicitly iterable, like an array or slice.This brings me to a point which is tangential to the main thrust of this comment, yet necessary for what follows: ranged integers should be defined inclusively, as closed intervals. I feel that this accords with how we reckon with integers already: I think of a
u8as[0, 255], not[0, 256): everything which fits in eight bits, nothing which does not.I do understand why using the exclusive range felt more natural, of course, but we really want it to be inclusive: most of this comment serves as the argument for that, although this point is not why I came to this issue today.
That would make
u0 == @Int(0,0),u1 == @Int(0,1), and so on. It wouldn't be possible to definenoreturnas an integer type, which I think is correct: an uninhabited integer range is not a necessary or especially coherent thing to have, the empty set of integers (if needed) can bevoid. We don't currently have a way to express an uninhabited integer type, and I don't see a compelling need for it: it could be spelled@Int(1,0), there's precedent, but I would vote against this. Zig might benefit from an uninhabited bottom type, whichnoreturnis not (#6602-comment ), but this shouldn't involve integers in any way.This is how Ada does it, and with reason. Ada also has a 'modular' type, which also specifies wraparound arithmetic for the type. Since this uses modulus, it mentions the first unsigned integer which isn't in the range. Zig uses explicit modular operators, so this is probably not a useful concept for us.
Another data point is the C standard, Section 7.20.1. It doesn't use interval notation at all, but it defines macros for limits in terms of minimum and maximum (inclusive) values.
Inclusive is the better convention in its own right, and also best compatible with the
forloop coercion just described. It would be weird if coercingu8to a range didn't include 255. That was an unexamined problem with #17039, writing something likefor (0..256) |byte: u8|is.. legible, but unsatisfyingly inconsistent, normally you can't say256andu8in the same breath. This is in fact a variation on why integer types should be described inclusive of their extreme values.I broadly agree with not having two range syntaxes for a
forloop, certainly not if they vary by one., and given that, "slice style" instead of "switch style" is the better choice. But both of these are useful nonetheless, and this would give us both.There's another place this coercion could apply: switch statements.
fn byteSwitch(b: u8) void { const a_range: type = @Int(0, 0x7f) const b_lead: type = @int(0xc2, 0xdf) // ... switch (b) { a_range => ... , b_range => ... , } }
These custom types coerce to
u8, and it should be perfectly clear what's going on here. Switch ranges are already inclusive, as they should be, and as integer ranges should be as well.For the same reason that it would be confusing to have an exclusive range in a switch statement, it would be confusing to have an exclusive range in a ranged integer definition.
Part of why I wrote this out was to be able to link to it later, but also on the premise that it's good to have discussions of the positive (or negative) consequences of a feature in the issue tracking it. I don't see a lot of point in writing the possible coercion features up as a separate issue until or unless this one is tagged 'accepted', although I'm willing to do so.
Reacted by Andrew Dunbar, Guillaume Wenzek and Axiomatic-MindI also wanted to respond separately to @SpexGuy's comment, since it has coherent implementation concerns which haven't yet been completely addressed. That was four years ago, and Spex might have a completely different take now, I want to acknowledge that before continuing. Zig has also changed quite a bit in that period of time.
For me, the benefit of this proposal is lifting ranges up into the type system, and the niche optimizations are a nice bonus. I think we can come up with rules which are general, easy to understand, and cover all cases of optional integer types, whether by themselves or in a packed struct.
At the limit, someone concerned if they left a niche in their enum type can use a test:
try expectEqual(@bitSizeOf(?E), @bitSizeOf(E));. I don't think losing an existing niche is that different from expanding the backing width of an enum which doesn't specify its integer type: they're both problems where it will either show up (invalidating the width of apacked struct) or if it's important it's worth a test. The literal first Zig file I ever wrote had a test asserting the size of the tagged union I was writing, I think it's basically ok to require users to assert invariants which need to stay that way.Another solution is this:
const ThreeEnum = enum(@Int(1,3)) { fee, fie, foe, };
This seems like something which would have to be legal if integer ranges are just integers. This guarantees (according to what follows) an optional sentinel at zero.
It seems to me that the solution to the limited-range problem of assigning comptime ints is something we'd have to do already: you need to provide a type for a
var, but not aconst. So it would be fine thatconst a = 4;gets the type@Int(4,4): but avarneeds more than a value, it needs a restriction on the future values it can take, and a comptime int value can't provide that. I'd suggest this would apply with or without integer range types, unless we wanted to punt and just say that acomptime_intcoerces to ausizeorisizewhen no type is provided. I would prefer that not happen.Niche Optimization of Ranged Integer Types
I propose a rule like the following: A ranged integer type has an automatic bitwidth, which is the minimum necessary to hold every value in the range. If there are any negative values in the range, this is a signed width, otherwise, unsigned.
If there are representable values in that width which are not part of the range, then the first value higher than the maximum, if it exists, is the optional sentinel, otherwise, the first value lower than the minimum.
If there are no representable values in the bitwidth (these are the integers we have already) then the original bitwidth is sign-extended by one bit, and the first value of the new width which is higher than the maximum of the old width is the optional.
I think this covers all the bases. Part of what Spex was addressing is combinatoric difficulties with the
align(0)part of the proposal, which were predicated on another proposal which didn't make it. This rule makes it simple enough to add integer ranges to packed structs, and it reflects the reality of the CPUs the code will ultimately run on.This rules out some of the more exotic 'optimizations', which would be space-optimizing only, and at the expense of everything else. For example, [-3, 240] could be represented in eight bits, this would require it to have a bitwidth of
i9instead. I think we actually want this property, just on the grounds of sanity: imagine debugging a [450, 700] which is stored inside a single byte! This is something we could do, but really, truly, shouldn't. That range would have au10bitwidth, and what we get in return is efficiency and coherence.Even in the event of following in the footsteps of Pascal and Ada by having arrays with widths and indices defined by a ranged integer, the offsetting of an index into that array should be at the point of use, not in the storage value of a ranged integer used for indexing that array. This is the tip of a huge iceberg, which I will now veer away from.
Alternative version of the rule which is actually better: for an unsigned range, if there are available values higher than the maximum, the highest value which fits in the bitwidth is chosen, otherwise a bit is added and the sentinel is the maximum value of the new width. For a signed range, if there are available values lower than the minimum, the lowest possible value of the width is chosen, otherwise the highest possible value, otherwise it is sign-extended by one and the lowest value of the new bitwidth is the optional sentinel.
My little bait-and-switch here is because that formula is significantly harder to understand, and looks worse at a first glance. But it's better, because it always chooses all 1 bits when it can, and when it can't, it chooses 0 followed by all 1s. This is the opposite of
?*T, but the same as?*Twill frequently not be available, and it's the exact opposite, which is nice. This is less for the compiler to think about, collapsing cases is generally a good thing.There might be a case for adding "if zero is not in the range, then zero is the sentinel" but I wouldn't make it. Someone else might, but I'm sensitive to the fact that the second version of the rule is more complex, and therefore harder to understand. I'll acknowledge that zero is a good value to be working with when it comes time to emit machine code, however.
I could also be wrong about the more complex rule being better for the compiler, but it would be easier to debug, and that's not nothing.
This neglects the actual size which the integer will take, and that's intentional. A
u9would have an optional bitwidth ofu10, and a sentinel of0b1111_1111_11, and when packed, would be exactly that many bits. On its own, or in an array, it would be of bitsize 16, or sizeof 2, and the sentinel would more than likely be0b1111_1111_1111_1111, although note that I'm assuming a convention which might not turn out to be the correct one, this is for illustration purposes only.As a note, Zig's documentation doesn't specify how non-natural-width integers are represented on the machine, and perhaps that's intentional, but, maybe it should.
Back to the comment I'm reply to: I don't think we want either of "bring your own bitwidth" or "bring your own sentinel". That's creating opportunities for incompatibility where there otherwise would not be. It seems to me that even in the packed case, a width wider than the minimum dictated by the rules here can be emulated by adding a pad field after the ranged integer.
It also brings the weirdness that a
u32[0,777]should coerce to au16[0,1000], since it inherently fits. But we don't let au32coerce to au16. This could be made to work, sure, but I think it's causing more problems than it solves.The major reason to even define what the sentinel will be is debugging, it isn't a value that user code has native access to, although presumably this would reveal it:
const opt_ranged: ?@Int(1, 14) = null; const as_cast: u8 = @bitCast(opt_ranged);
And I think the fact that the above would be possible, and debugging, are good enough reasons to have the sentinel value be a specified part of the language.
This may be irrelevant, given that
align(0)is closed, but I also see little value in allowing enums to combine in thealign(0)style of the first post. One consequence of that would be making hex dumps and other debug info very opaque indeed. Set unions of enums have been proposed before, and they have the same basic issues as densely packing logically-separate enums into a single structure: instead of one concrete value, enums could have up to several, and that is most likely the wrong tradeoff to make.I'd argue that even for packed structs, trying to minimize space in a range like [256, 511] should be resisted. If it's imperative that values of that nature fit in eight bits and not nine, it's reasonable to do that with arithmetic. Things would just get weird otherwise.
The partitioned integer is a cool idea though. I haven't a clue how feasible it is, but it's probably its own issue rather than this one.
You'll have to forgive that all of my example code here uses C++ syntax, I'm not familiar enough with Zig to be confident giving the Zig equivalent.
My C++ library defines the range bounds inclusively (
integer<0, 3>can take any of the values in the set{0, 1, 2, 3}), and in my usage that has almost always been the more convenient choice and I can't remember any examples where it was confusing and caused a bug. The primary time that comes to mind where I wanted a half-open range where when I want to get the index type of a range (integer<0, max_size - 1>), but I just added anindex_typethat does the that for me and then never worried about it again.I agree that we should make the representation of the integer type match its value. In other words,
integer<10, 15>(10)should store bits that represent10, not0. This means that an integer type likeinteger<256, 260>requires 2 bytes of storage even though it theoretically only needs 1, but in exchange arithmetic and comparisons are a single instruction. In cases where you are very size constrained, it's fairly easy to write apacked_integertype that performs the necessary arithmetic adjustments for you, but you'll get much better code generation for the common case by not offsetting.My integer type also works with a compressed optional implementation -- it's split into the general optional type and an associated traits class to manage the optimization for integer types specifically. The general idea is similar to what was outlined here: use the unused values as sentinel values. My implementation uses the least representable value as the first tombstone value and counts up from there, skipping over the valid value range. So if
integer<1, 10>is represented by auint8_t, a disengaged optional of that type would be represented by0x00. Thetombstone_traitsidea allows space-efficientoptional<optional<T>>, as well as more general tombstone-style code (for instance, some hash maps need two tombstone values, so in that example the two representations would be0x00and0x0A). I chose this only because that ensures that for unsigned underlying storage, the representation of a disengagedoptional<T>is all0x00bytes. I don't know if that actually matters for anything, but every other choice seemed more complex or even less likely to have a beneficial representation.I much prefer the design of my library where the user says "I want an integer in this range, you figure out how to store it (
integer<5, 10>)" over a design where the user says "I want an integer in this range that is stored in this underlying type (u8<5, 10>)".Reacted by Sam Atman, Guillaume Wenzek and Axiomatic-MindIt's probably not clear by casual inspection and I don't know how optional support works in Zig, but in my C++ library, the type
optional<integer<0, 254>>takes 1 byte of storage with the value of255being the sentinel and it follows the compressed code path, but the typeoptional<integer<0, 255>>takes two bytes of storage and it follows the regular uncompressed codepath that is used by default for any types that don't have the ability to store sentinel values in them. This is shown inoptional.cppby looking at theoptional<nullable T>specialization vs theoptional<typename T>base template.Reacted by Sam Atman and Guillaume WenzekJust chiming in that there's also a nice interaction between this and the "named integer types" shown in https://ziglang.org/devlog/2024/#2024-11-04 for defining named types with massive but specific ranges of valid values:
pub const SpecialType = enum(@Int(40, 40_000_000)) { _, };Would this solve the following codegen issue, or should I open a new issue?
In the following
getPtrfunction, Zig uses au64forhi_bitsandtag_bitswhere it could use smaller types transparently without changing the semantics.const assert = @import("std").debug.assert; const max = @import("std").math.maxInt; /// Double or NaN-packed value. const Value = packed union { double: f64, bits: u64 }; /// Check if value is a non-null packed pointer with the given tag; return null on failure. export fn getPtr(v: Value, tag: u8) ?*anyopaque { assert(tag < 8); const ptr_val = (v.bits & (max(u45) << 3)); // Actual address const hi_bits = (v.bits >> 51); // 13 high bits identifying a pointer const tag_bits = (v.bits & 7); // Pointer-destination type tag const is_ptr = hi_bits == 0b0111111111111; const is_tag = tag_bits == tag; return if (is_ptr and is_tag) @ptrFromInt(ptr_val) else null; } // Same but with better codegen thanks to explicit types. export fn getPtr2(v: Value, tag: u8) ?*anyopaque { assert(tag < 8); const ptr_val: u48 = @intCast(v.bits & (max(u45) << 3)); const hi_bits: u13 = @intCast(v.bits >> 51); // BTW: Same codegen with u16 here const tag_bits: u3 = @intCast(v.bits & 7); // BTW: Same codegen with u8 here const is_ptr = hi_bits == 0b0111111111111; const is_tag = tag_bits == tag; return if (is_ptr and is_tag) @ptrFromInt(ptr_val) else null; }
It should be clear from static analysis that
hi_bitscould be safely stored in au16even though it's inferred as au64. This is because the value it receives is certain to only have some of its lowest 13 bits set, and is only ever used for comparison to a value with its lowest 12 bits set.Likewise,
tag_bitscould be stored in au8without changing the semantics, since it receives a value that's certain to use at most 3 bits, and it's only used for comparison with au8(even asserted to fit in au3).Despite this, the codegen for the first implementation jumps through hoops to preserve the full 64 bits of
hi_bitsandtag_bitsfrom what I can tell; I'm not good at reading assembly yet but it's clear that it's a bit worse:getPtr1: movabs rcx, 281474976710648 and rcx, rdi movabs rax, -2251799813685248 and rax, rdi and edi, 7 movabs rdx, 9221120237041090560 xor rdx, rax movzx esi, sil xor rsi, rdi xor eax, eax or rsi, rdx cmove rax, rcx ret getPtr2: movabs rax, 281474976710648 and rax, rdi mov rcx, rdi shr rcx, 51 and dil, 7 xor edx, edx cmp dil, sil cmovne rax, rdx cmp ecx, 4095 cmovne rax, rdx ret
Tell me if I'm expecting too much of static analysis. :-)
@mnemnion mentioned some objections to the upper bound being exclusive. In addition, another issue came up on discord: if this proposal is combined with #15737 it would result in some strange/unintuitive resolutions. For example, adding two
@Int(0,10)s would resolve to@Int(0, 19), and multiplying two@Int(0, 10)s would resolve to@Int(0, 82).Reacted by ominitay, Sam Atman and Guillaume WenzekI have a few other points against the upper bound being exclusive:
-
It makes dealing with integers representing an index into an array more tricky. With the current proposal, I feel myself wanting to type
@Int(0, array.len)to represent indexes into the array. If we could only represent valid indexes into an array, loops like:while (idx < array.len) : (idx += 1) { ... }won't work, asidxcould not representarray.lenat the end of the loop. While we could do@Int(0, array.len + 1)orwhile (true) : (idx += 1) { ... if (idx == array.len) break; }, neither are intuitive. EDIT: A for loop could also work here, where we cast the loop index to the array index type, but many people would likely just use the loop index at that point, bypassing further checks that could be done with indexing math inside the loop. EDIT 2: With sentinel terminated slices, you would need to access the index atarray.lenanyways, so my original point still stands here. -
Generic code might be a bit harder, as the upper bound is never guaranteed to be represented within the type that
@Int()may return. A side effect is that au65535may never be constructed without the usage of acomptime_int. Perhaps this point is niche, but I thought it worth mentioning as well. -
If the upper bound is exclusive, the compiler needs to special case
@Int(0, 0)asunreachableor something similar, whereas it would simply be au0if the upper bound were inclusive.
Reacted by drfuchs, Sam Atman, 00JCIV00, DivergentClouds, Elijah M. Immer, Axiomatic-Mind and Rue-
- added 3 commits that reference this issue
on May 17, 2025 - added a commit that references this issue
on Aug 4, 2025 - marked Add erroring and widening arithmetic operators #14320 as a duplicate of this issue
on Apr 24, 2026
Zig already has ranged integer types, however the range is required to be a signed or unsigned power of 2. This proposal is for generalizing them further, to allow any arbitrary range.
Let's consider some reasons to do this:
One common practice for C developers is to use
-1orMAX_UINT32(and related) constants as an in-bound indicator of metadata. For example, the stage1 compiler uses asize_tfield to indicate the ABI size of a type, but the valueSIZE_MAXis used to indicate that the size is not yet computed.In Zig we want people to use Optionals for this, but there's a catch: the in-bound special value uses less memory for the type. In Zig on 64-bit targets,
@sizeOf(usize) == 8and@sizeOf(?usize) == 16. That's a huge cost to pay, for something that could take up 0 bits of information if you are willing to give up a single value inside the range of ausize.With ranged integers, this could be made type-safe:
Now, not only do we have the Optionals feature of zig protecting against accidentally using a very large integer when it is supposed to indicate
null, but we also have the compile error helping out with range checks. One can choose to deal with the larger ranged value by handling the possibility, and returning an error, or with@intCast, which inserts a handy safety check.How about if there are 2 special values rather than 1?
Here, size of N would be 4 bytes.
Let's consider another example, with enums.
Enums allow defining a set of possible values for a type:
There are 3 possible values of this type, so Zig chooses to use
u2for the tag type. It will require 1 byte to represent it, wasting 6 bits. If you wrap it in an optional, that will be 16 bits to represent something that, according to information theory, requires only 2 bits. And Zig's hands are tied; because currently each field requires ABI alignment, each byte is necessary.If #3802 is accepted and implemented, and the
is_nullbit of optionals becomesalign(0), then?Ecan remain 1 byte, and?Ein a struct withalign(0)will take up 3 bits.However, consider if the enum was allowed to choose a ranged integer type. It would choose
@Int(0, 3). Wrapped in an optional, it actually could choose to use the integer value 3 as theis_nullbit. Then?Ein a struct will take up 2 bits.Again, assuming #3802 is implemented, Zig would even be able to "flatten" several enums into the same integer:
If you add up all the bits of the tag type, it comes out to 10, meaning that the size of MyFancyType would have to be 2 bytes. However, with ranged integers as tag types, zig would be able to flatten out all the enum tag values into one byte. In fact there are only 21 total tag types here, leaving room for 235 more total tags before MyFancyType would have to gain another byte of size.
This proposal would solve #747. Peer type resolution of comptime ints would produce a ranged integer:
With optional pointers, Zig has an optimization to use the zero address as the null value. The
allowzeroproperty can be used to indicate that the address 0 is valid. This is effectively treating the address as a ranged integer type! This optimization for optional pointers could now be described in userland types:One possible extension to this proposal would be to allow pointer types to override the address integer type. Rather than
allowzerowhich is single purpose, they could do something like this:This would also introduce type-safety to using more than just 0x0 as a special pointer address, which is perfectly acceptable on most hosted operating systems, and also typically set up in freestanding environments as well. Typically, the entire first page of memory is unmapped, and often the virtual address space is limited to 48 bits making
@Int(os.page_size, 1 << 48)a good default address type for pointers on many targets! Combining this with the fact that pointers also have alignment bits to play with, this would give Zig's type system the ability to pack a lot of data into pointers which are annotated withalign(0).What about two's complement wrapping math operations? Two's complement only works on powers-of-two integer types. Wrapping math operations would not be allowed on non-power-of-two integer types. Compile error.