[{"content":"I recently looked at P4313R1, a standards proposal paper which adds a set of bitmask operations for enums, using a C++26 annotation to opt-in. The core idea being that, to use an example from the paper, given code like the below:\nenum class [[=std::bitmask_type]] Permission { None = 0, Read = 1 \u0026lt;\u0026lt; 0, Write = 1 \u0026lt;\u0026lt; 1, Execute = 1 \u0026lt;\u0026lt; 2, }; The [[=std::bitmask_type]] annotation would automatically imbue Permission with an accessible set of bitwise operations so that you as a user needn\u0026rsquo;t write them out yourself. This got me thinking about the murky underbelly of C++ integer operations. Your C++ compiler will happily compute bitwise operations for any integral type as a builtin operation. This extends to types which we don\u0026rsquo;t traditionally think of as integers, such as wchar_t, the UTF character types char8_t through char32_t, and bool. But slightly more happens here than meets the eye. Because when you attempt to perform an operation like a | b, and the type of a and b is an integral type smaller than int, the language does not operate on the bit patterns of a and b directly. They undergo integral promotion - they are promoted up to int as an intermediate state, the bit patterns of these two ints are combined, and the result is returned to you, still as an int.1 Consider the below:\n//Two shorts constexpr short perm_A {1 \u0026lt;\u0026lt; 0}; constexpr short perm_B {1 \u0026lt;\u0026lt; 1}; //And the result type of running a bitwise operation on them is int static_assert(std::same_as\u0026lt;decltype(perm_A | perm_B), int\u0026gt;); Now, for the most part this is harmless - if you cast the above perm_A | perm_B back to short then the unnecessary bytes are truncated away and you\u0026rsquo;re left with a short which contains exactly the value it would have held if you\u0026rsquo;d combined perm_A and perm_B as short directly. And if you want to be certain that you are keeping your types consistent, and avoid the pernicious bugs of silent narrowing conversions, you can get into the habit of static_cast-ing the result of your bitwise operation to its original type. The important thing here however is that integral promotion is not optional. Unlike most other areas of C++ the developer doesn\u0026rsquo;t get a choice - your types will unavoidably be promoted for these operations.\nThe second part of what makes this a hazard for bool specifically is boolean conversion, a special case which does not truncate but instead explicitly converts zero to false and non-zero values to true. For these non-zero values, whatever bit pattern was previously stored is discarded and replaced with true, which has an integer value of 1. So if we take code like this:\nint x{10}; bool b{static_cast\u0026lt;bool\u0026gt;(x)}; and look at the generated asm (in this case from unoptimised x86-64 gcc 16.2):\nmov DWORD PTR [rbp-4], 10 ;Store the value of 10 cmp DWORD PTR [rbp-4], 0 ;Then compare to 0, set ZF if x is zero setne al ;Write 1 to AL if ZF is clear mov BYTE PTR [rbp-5], al ;Store the result The standard (specifically [conv.bool]) bases converting to bool on the only possible values which bool can hold - true and false, regardless of whatever bit pattern may have originally been used to create them.\nPutting it together With all of that covered, let\u0026rsquo;s talk through what happens when you try to evaluate ~true and then cast the result to a bool:\nThe value true is promoted to an int with a value of 1. The complement of 1 as an int is calculated as -2, as two\u0026rsquo;s complement behaviour is required as of C++20. The value of -2 is then cast back down to bool, undergoes boolean conversion, and since -2 is non-zero, becomes true. There you have it, the complement of true is true, or spelled in C++ static_cast\u0026lt;bool\u0026gt;(~true) == true.\nThis brings us back to enums. An enum is permitted to use any integral type as its underlying type, including our good friend bool. So let\u0026rsquo;s define one:\nenum class [[=std::bitmask_type]] boolean : bool{ FALSE, TRUE, }; Let\u0026rsquo;s also look at the bitwise complement operator as laid out in P4313R1:\ntemplate\u0026lt;bitmask-like T\u0026gt; constexpr T operator~ (T lhs) noexcept { return static_cast\u0026lt;T\u0026gt;(~to_underlying(lhs)); } By now the workings of that operator should be a familiar shape - first we convert the enum to its underlying type (in our case bool), then integral promotion is applied if that type is lower ranked than int, then we perform the bitwise operation, then we cast back to the enum type. As we would expect:\nstatic_assert(~boolean::TRUE == boolean::TRUE); Godbolt here.\nThis has all the potential to be a slightly confusing corner case in the language.\nWhen it\u0026rsquo;s false There is one exceptional case here which compounds the problem. If we try to run that same example in gcc, we get a different result:\nstatic_assert(~boolean::TRUE == boolean::FALSE); Godbolt here.\nSo what is going on with gcc? If we move the operation to runtime by defining two functions which depend on it as a runtime value:\nenum class boolean : bool{ FALSE, TRUE, }; void f(int x, bool\u0026amp; out) { out = static_cast\u0026lt;bool\u0026gt;(x); } void g(int x, boolean\u0026amp; out) { out = static_cast\u0026lt;boolean\u0026gt;(x); } and look at the asm (again unoptimised x86-64), then we get:\n\u0026#34;f(int, bool\u0026amp;)\u0026#34;: push rbp mov rbp, rsp mov DWORD PTR [rbp-4], edi mov QWORD PTR [rbp-16], rsi cmp DWORD PTR [rbp-4], 0 setne dl ;same setne pattern from [conv.bool] earlier mov rax, QWORD PTR [rbp-16] mov BYTE PTR [rax], dl nop pop rbp ret \u0026#34;g(int, boolean\u0026amp;)\u0026#34;: push rbp mov rbp, rsp mov DWORD PTR [rbp-4], edi mov QWORD PTR [rbp-16], rsi mov eax, DWORD PTR [rbp-4] and eax, 1 ;store the low bit in eax mov rdx, QWORD PTR [rbp-16] mov BYTE PTR [rdx], al ;and move al to the out-param nop pop rbp ret What we see is that when the type is not spelled bool, gcc will truncate down to the lowest bit rather than perform a boolean conversion. Any even value, when cast down, will produce boolean::FALSE, and any odd one will produce boolean::TRUE.\nThis is unique to enum types specifically, gcc generally performs consistently when handling bool directly:\nconstexpr boolean b{boolean::TRUE}; static_assert(static_cast\u0026lt;bool\u0026gt;(~b) == false); static_assert(static_cast\u0026lt;bool\u0026gt;(~true) == true); static_assert(std::to_underlying(~b) == false); static_assert(static_cast\u0026lt;bool\u0026gt;(~std::to_underlying(b)) == true); In contrast, Clang and MSVC will evaluate the result of all four of these complement-and-cast combinations as true, which is what the standard says they must be. Full godbolt comparison here. This seems to be a fairly long-lived conformance bug in gcc.\nUnscoped enums, promotion, and undefined behaviour But there is one place in the standard where integral promotion rules go further and introduce a new vector for undefined behaviour to enter into your program when doing these bitmask operations. Consider this code:\nenum nums{ zero, one, two, three, }; //This is implementation defined but holds on gcc and Clang. static_assert(std::is_same_v\u0026lt;std::underlying_type_t\u0026lt;nums\u0026gt;, unsigned int\u0026gt;); //nums uses an unsigned int as its base, therefore the entire domain of representable values is \u0026gt;= 0. //So let\u0026#39;s test this: static_assert(~one \u0026gt;= 0); //FAILS Godbolt here\nLooking at the diagnostic, gcc gives us:\n\u0026lt;source\u0026gt;:15:20: error: static assertion failed 15 | static_assert(~one \u0026gt;= 0); //FAILS | ~~~~~^~~~ • the comparison reduces to \u0026#39;(-2 \u0026gt;= 0)\u0026#39; So what happened here? Well, the passage in the standard which covers integral promotion, [conv.prom], carves out a special bullet for unscoped enumeration types with no fixed underlying type to behave differently from integer types when promoting. If the entire range of values can be stored in an int, then that is the type which it promotes to, then it tries unsigned int, and if that fails it repeats this signed-then-unsigned pattern for long and long long until it finds a suitable type. But, where this differs from plain integer types is that it only considers the effective range of representable values; and so an enum backed by unsigned int will promote to int regardless of the normal conversion ranking which would forbid it for integral types. As before, the fact that you are calling the builtin operator via ~one will unavoidably promote it to int, its complement is calculated as -2, and then the comparison is performed between two ints, and fails. If instead you convert to the underlying type first then you really see the asymmetry - static_assert(~std::to_underlying(one) \u0026gt;= 0); succeeds, and is so vacuously true that gcc even warns about its redundancy.\nBut this is only half the battle, and we need to talk about converting the result of our bitwise operation back to the enum type, and this is where UB creeps in. An unscoped enum with no specified underlying type defines itself in terms of a valid range of values determined by its enumerators independently of the range of whatever actual type is used by the compiler to back it. [dcl.enum] tells us the value range of such an enum is the value range of a hypothetical integer type of width M, where M is the minimum bits required to represent all enumerators. To demonstrate, consider a few examples:\nenum small{ //Range of enumerators is 0..1, so value range is that of an unsigned, one-bit integer. a = 0, b = 1, }; enum medium{ //Range of enumerators is 0..7, so value range is that of an unsigned, three-bit integer c = 0, d = 4, e = 7 }; enum large{ //Range of enumerators is -1..100, so value range is that of a signed, eight-bit integer f = -1, g = 10, h = 100, }; Even though your compiler will quite likely store these enums as unsigned int and int, only a subset of possible representable values is valid to use - that of the M-width integer. [expr.static.cast] in the standard makes it clear that it is undefined behaviour to cast a value outside of that range to the enum type. So, using the enumerators of the above example:\nstatic_cast\u0026lt;small\u0026gt;(1); //Well-defined static_cast\u0026lt;small\u0026gt;(2); //UB static_cast\u0026lt;medium\u0026gt;(3); //Well-defined: while not a value of an enumerator, it is within the range of an unsigned, three-bit integer static_cast\u0026lt;medium\u0026gt;(8); //UB static_cast\u0026lt;large\u0026gt;(-128); //Well-defined static_cast\u0026lt;large\u0026gt;(128); //UB This is our danger zone. The bitwise operators, in particular operator~, can produce values which are out of the well-defined value range for an enum, and user-defined operators which attempt to cast this value back down to the enum type can invoke UB.\nBut, I hear you cry, what about the example earlier with ~one? We saw that give a value of -2 which is outside the value range of nums. This is true, but here is where integral promotion actually saves us - one was promoted to int before we calculated its complement, and it was never cast back down again to invoke our UB. We only get into dangerous territory with a user-defined operator which converts the result of the operation back down to the original enum type.\nIt is also important to note that this UB risk only applies to unscoped \u0026ldquo;plain\u0026rdquo; enums with no fixed underlying type. All scoped enum types have a fixed underlying type, even if they don\u0026rsquo;t specify it (in that case, int); and all unscoped enum types which do specify a fixed underlying type inherit that type\u0026rsquo;s value range instead of inventing their own.\nAvoiding this in your own code Let\u0026rsquo;s say that, like P4313, you are adding bitmask operations to enums for your own code. Now that we have more weirdness in this area of C++, how would you go about making sure that your code is correct and unconfusing? Avoiding the bool case is simple - just constrain away that the underlying type can\u0026rsquo;t be bool, either for all enums or for the complement operator. But the standard doesn\u0026rsquo;t come with a handy type trait or reflection metafunction to specifically detect an unscoped enum with no fixed underlying type. Fortunately we can make one:\ntemplate\u0026lt;typename E\u0026gt; concept unfixed_enum = std::is_enum_v\u0026lt;E\u0026gt; \u0026amp;\u0026amp; !requires { E{0}; }; [dcl.init.list] only permits this initialization from a scalar for enums with a fixed underlying type, and 0 is the only possible value which is representable for all possible enums. To cover the edge cases, the standard defines an enum with no enumerators as having the value range of an unsigned, one-bit integer; and the thing which excludes 1 as a possible value is enum E{ a = -1 };, which has a range [-1, 0]. Be sure not to forget the std::is_enum_v. After all, there are many types for which T{0} is ill-formed, and most of them are not enums.\nThis post has mostly been focused on the bitwise complement operator, but the UB trap of an unfixed enum also applies to the bitwise left shift operator. If you take the route of constraining per operation, don\u0026rsquo;t let it slip your mind.\nThis brings us back once again to P4313R1. At time of writing, the concept which constrains these operators only constrains the type to be some enumeration type annotated with [[=std::bitmask_type]]. As such, it will allow confusing behaviour on enum types which are backed by bool, and potential UB on plain enums with no fixed underlying type. If the authors don\u0026rsquo;t want to standardise the above init-list trick, they could constrain it on std::is_scoped_enum as a conservative approach to prevent users from having easy access to the UB, since all unscoped enums can use the builtin bitwise operators anyway (albeit returning int). They could also constrain the underlying type to not be bool to minimise the confusion; particularly as the diagnostic which would normally get issued for complementing a bool in user code would be suppressed by default if it came from a system header. I do intend to contact the authors of P4313 about this to see if they want to add these constraints to their paper.\nAside: What about floating point? You might be wondering how the standard defines conversions of floating point numbers to these bool-backed enums. It is perhaps unsurprising - [expr.static.cast] requires that the behaviour is equivalent to first converting the floating point value to the underlying type of the enum, and then converting it to the enum type; pointing the user to [conv.fpint] which has an explicit note directing readers to our old friend [conv.bool]. Per the standard, this should mean that any floating point value other than exactly 0.0 or -0.0 would convert to boolean::TRUE, otherwise we end up back at boolean::FALSE.\nSo let\u0026rsquo;s start small:\nenum class boolean : bool { FALSE, TRUE, }; static_assert(static_cast\u0026lt;boolean\u0026gt;(0.0) == boolean::FALSE); static_assert(static_cast\u0026lt;boolean\u0026gt;(0.5) == boolean::TRUE); static_assert(static_cast\u0026lt;boolean\u0026gt;(2.5) == boolean::TRUE); Godbolt here.\nBoth Clang and gcc reject this code. The errors are that (boolean)2.5e+0 is not a constant expression; and static_cast\u0026lt;boolean\u0026gt;(0.5) == boolean::TRUE is wrong, and is instead equal to boolean::FALSE. MSVC accepts the above code.\nBut let\u0026rsquo;s go further and inspect what actually gets stored there. We write a function to examine the bit pattern generated from these casts, then try some values:\nvoid show(double d) { const boolean e = static_cast\u0026lt;boolean\u0026gt;(d); unsigned char byte {}; std::memcpy(\u0026amp;byte, \u0026amp;e, 1); std::println(\u0026#34;static_cast\u0026lt;boolean\u0026gt;({:7.1f}) stored byte {:3} required {}\u0026#34;, d, byte, (d != 0.0) ? 1 : 0); } int main() { show(0.0); show(0.5); show(1.0); show(2.5); show(3.0); show(256.0); show(-2.5); } Godbolt here.\nThe above code compiles without warning on all three compilers, but while MSVC again does the right thing in all cases, gcc and Clang\u0026rsquo;s behaviour is much more worrisome. Looking at the output, we see that gcc will always just truncate the value to an integer, then load that bit pattern into the resulting boolean. So 2.5 truncates to 2 and gives a boolean whose underlying bit pattern is 2. Clang does the same thing on unoptimised builds, but when optimisation is turned on, the truncated value then goes through the [conv.bool] transformation as normal and produces the right answer for all non-zero values outside of the range (-1, 1).\nThis is more concerning than some funky bit patterns, however. The valid range of boolean is [0, 1]. It is UB to read objects with bit patterns outside of this range. But, I hear you ask, what happens if we take these booleans and cast them back to bool, triggering a boolean conversion with no initial floating point state? MSVC again does the right thing; gcc doesn\u0026rsquo;t modify the bit patterns, leaving us with bool which are out of range of the type (and therefore UB to read); and Clang is where the fun happens again. Running the above code with a cast from e to a bool, and then memcpy-ing that bool into the unsigned char, we get this result for unoptimised builds:\n0.0 0.5 1.0 2.5 3.0 256.0 -2.5 required 0 1 1 1 1 1 1 gcc 14 / 15 / 16 0 0 1 2 3 0 254 clang 19 / 20 / 21 0 0 1 0 1 0 0 clang 23.1.1 0 0 1 1 1 0 1 Versions of Clang before 23 appear to simply perform trunc(d) \u0026amp; 1; meaning that after truncation, odd numbers are true and even numbers are false. Clang 23 performs (trunc(d) % 256) != 0, narrowing to a byte first. So 256.0 becomes FALSE and 257.0 becomes TRUE. Optimised builds act as before - truncating then doing the proper boolean conversion.\nUltimately this all comes to a rather absurd head. Consider the below code:\n#include \u0026lt;print\u0026gt; enum class boolean : bool { FALSE, TRUE, }; //Force this to be runtime boolean make(double d) { return static_cast\u0026lt;boolean\u0026gt;(d); } int main() { const boolean e = make(2.5); std::println(\u0026#34;e == boolean::TRUE : {}\u0026#34;, e == boolean::TRUE); std::println(\u0026#34;e == boolean::FALSE : {}\u0026#34;, e == boolean::FALSE); const bool b = static_cast\u0026lt;bool\u0026gt;(e); std::println(\u0026#34;b == true : {}\u0026#34;, b == true); std::println(\u0026#34;b == false : {}\u0026#34;, b == false); std::println(\u0026#34;b + 0 : {}\u0026#34;, b + 0); } Godbolt here.\nMSVC again leads the pack in giving the correct answer. Clang gets it exactly wrong and thinks that e is FALSE and b is false. gcc takes things up a notch, by providing an instance of an enum which compares equal to all of its enumerators, and a bool which is neither true nor false, until you turn on the optimiser and get a bool which is both true and false.\nAnd to hit all the obligatory targets of any conversation on floating point, infinity and NaN show as 0 and become FALSE as boolean and so cast to false as bool on gcc and Clang; and give 1 and therefore TRUE and true on MSVC.\nIn lighter news, Clang trunk seems to have fixed the above example to behave correctly, so some of the issues described in this section should be patched out soon.\nConclusion What did we learn from all this? Perhaps that even sensible and uncontentious pure-library-level papers such as P4313 should be on their guard against a footgun from some exotic C++ edge case. Or perhaps, in more practical terms:\nIntegral promotion is unavoidable. If you think you are operating on an integer which is smaller than int, there\u0026rsquo;s a good chance it became an int right under your nose. The bitwise complement of a bool is always true, even when wrapped in an enum. However you should not rely on this behaviour as gcc has a longstanding bug which gets this wrong (or right, if you prefer logic to C++). Unscoped enumeration types with no fixed underlying type have a range of valid values which may be smaller than that of whatever type actually underlies them. To be clear, the ranking order of integral promotion is more complex than just \u0026ldquo;convert to int\u0026rdquo;. For an integer type of lower conversion rank (not size) than int, if all values of that type can be represented as int then int is chosen. Otherwise unsigned int is chosen. In practice this will usually come out as int; but on implementations where int and short are the same size, unsigned short will promote directly to unsigned int. And for other types such as the charN_t family, conversion can proceed to longer integer types such as long and long long. bool is singled out as a special case and always promotes to int.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://dryperspective.github.io/posts/complement-of-true/","summary":"Consequences of integral promotion, UB on unscoped enums, and why boolean conversion is surprising.","title":"The complement of true is true, except when it's false"}]