Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
f3cc6c9
Tests: Add the integer-signedness C reference
r41k0u Sep 4, 2026
daf429d
Core: Add the signedness machinery (no behaviour change)
r41k0u Sep 4, 2026
4fa6fc5
Core: Route every integer widening and narrowing through convert()
r41k0u Sep 4, 2026
f266f1f
Core: Carry signedness from every declaration and inference site
r41k0u Sep 4, 2026
5f36bd2
Core: Type each binary operation per node with C's arithmetic convers…
r41k0u Sep 4, 2026
db12298
Core: Choose signed or unsigned operators and predicates from the pro…
r41k0u Sep 4, 2026
6ab992e
Core: Fold a negated integer literal into a constant
r41k0u Sep 4, 2026
568a651
Core: Type comparison and boolean results as unsigned i1
r41k0u Sep 4, 2026
4c65510
Tests: Integer signedness cases mirroring the C reference
r41k0u Sep 4, 2026
6f48498
Docs: Integer semantics user-guide page
r41k0u Sep 4, 2026
0273a67
Merge branch 'feat/global-variables' into feat/integer-signedness
r41k0u Sep 17, 2026
5a75c16
Merge remote-tracking branch 'origin/master' into feat/integer-signed…
r41k0u Sep 17, 2026
359ddbd
Merge remote-tracking branch 'origin/master' into feat/integer-signed…
r41k0u Sep 18, 2026
8e802e8
Core: Return through the same value path as every other consumer
r41k0u Sep 18, 2026
ca92f2a
Core: Rank an enum constant as C's int, not as i64
r41k0u Sep 18, 2026
eaaa12e
Tests: Return of a map value, return of a struct, enum-constant rank
r41k0u Sep 18, 2026
a230b44
Core: Give bools C's rules: widen to 0 or 1, narrow by comparing with…
r41k0u Sep 24, 2026
78ec75a
Tests: Bool widening and narrowing
r41k0u Sep 24, 2026
c52a5f5
Merge branch 'master' into feat/integer-signedness
r41k0u Sep 24, 2026
26604b9
Core: Give a helper call's value the sign its registry entry declares
r41k0u Sep 24, 2026
532dc0f
Core: Rank ctx fields and map values by their declared types
r41k0u Sep 24, 2026
9cd8e10
Core: Review nits on the signedness branch
r41k0u Sep 24, 2026
e35d4c0
Core: Pass the result sign to the operator in augmented assignment
r41k0u Sep 24, 2026
6b19bf6
Tests: Unsigned augmented assignment, helper results, map values, ctx…
r41k0u Sep 24, 2026
6caab6f
Tests: Rank the ctx-field case on a scalar field, not the packet pointer
r41k0u Sep 24, 2026
f727bfc
Docs: Note the packet-pointer-field gotcha, with the planned mechanism
r41k0u Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,7 @@ user-guide/maps
user-guide/structs
user-guide/compilation
user-guide/helpers
user-guide/integers
```

```{toctree}
Expand Down
3 changes: 3 additions & 0 deletions docs/user-guide/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,9 @@ PythonBPF uses Python's `ctypes` module for type definitions:
* `c_void_p` - Void pointers
* `str(N)` - Fixed-length strings (e.g., `str(16)` for 16-byte string)

Integers follow C's rules for width, sign, conversion and arithmetic; see
{doc}`integers` for the details and the places where this differs from Python.

## Example Structure

A typical PythonBPF program follows this structure:
Expand Down
153 changes: 153 additions & 0 deletions docs/user-guide/integers.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
# Integer Semantics

PythonBPF programs are Python syntax, but the integers in them behave as C integers: the
program runs in the kernel as BPF bytecode, where every value is a fixed-width machine
word. This page describes the rules the compiler applies. They are C's rules, applied to
the `ctypes` types you declare, so a program's arithmetic matches what the equivalent C
program compiled with clang would compute.

```{note}
This is one of the few places where PythonBPF deliberately differs from Python. Python
integers have arbitrary precision and no unsigned types; BPF has neither. The
[divergences from Python](#divergences-from-python) are listed at the end of this page.
```

## Types

An integer's type is the `ctypes` type it was declared with, and the type carries both
a width and a sign:

| Signed | Unsigned | Width |
|---|---|---|
| `c_int8` | `c_uint8` | 8 |
| `c_int16` | `c_uint16` | 16 |
| `c_int32` | `c_uint32` | 32 |
| `c_int64` | `c_uint64` | 64 |

Every declaration site uses these types: local variables initialised with a constructor
call, `@bpfglobal` variables, `@struct` fields, map keys and values, and fields read from
`vmlinux` structures. Helper functions return the type of the kernel's signature, so
`pid()` and `ktime()` are unsigned while `probe_read`-style helpers return a signed
`long`.

```python
count = c_uint32(0) # a 32-bit unsigned local
delta = c_int64(-1) # a 64-bit signed local
```

A local assigned without a constructor takes its type from the expression:

```python
now = ktime() # c_uint64, the helper's return type
total = count + 1 # the type of the addition (see below), held in a 64-bit slot
```

Undeclared locals are always 64 bits wide; the inferred type only decides their sign.
Declare the local with a constructor when a narrower width matters.

### Literals

A literal has the type a C compiler gives it: `int` (32-bit signed) if the value fits,
`long long` (64-bit signed) otherwise. This matters for mixed arithmetic: in
`count / -2` with `count` a `c_uint32`, the literal `-2` is a 32-bit `int`, so the
division happens in `c_uint32` exactly as it would in C.

## Assignment and conversion

Assigning a value to a variable of a different integer type converts it, and the
variable's declared type is what the stored value *is* afterwards:

* **Widening preserves the value.** The conversion looks at the *source*'s sign: an
unsigned source is zero-extended, a signed source is sign-extended. So a `c_uint32`
holding `0xFFFFFFFF` stored into a `c_int64` gives `4294967295`, and a `c_int32`
holding `-1` stored into a `c_uint64` gives `0xFFFFFFFFFFFFFFFF`. This is what C and
`ctypes` both do.
* **Narrowing truncates.** Only the low bits survive.
* **Same width reinterprets.** A `c_uint32` `0xFFFFFFFF` stored into a `c_int32` reads
as `-1`.

The same rules apply to explicit conversions written as constructor calls
(`c_int64(count)`), to struct field stores and to `return`.

## Arithmetic

Each binary operation is typed on its own, from its two operands, following C's usual
arithmetic conversions:

1. Operands narrower than 32 bits are promoted to `c_int32`.
2. If both operands have the same sign, the result has the wider width and that sign.
3. If the signs differ, the unsigned type wins when it is at least as wide as the signed
one; otherwise the signed type wins.

The operation is then performed in that type, and its result has that type. The variable
receiving the result plays no part until the final store. Two consequences worth knowing:

* **Intermediate results wrap at their own width.** `c_uint32(0x80000000) * c_uint32(2)`
is a `c_uint32` multiplication, so it wraps to `0` before being stored, even if the
destination is a `c_uint64`. Widen an operand first if you want a 64-bit product.
* **Mixed signs go unsigned.** `c_uint32(10) / c_int32(-2)` is an unsigned division by
`0xFFFFFFFE`, giving `0`, not `-5`.

The sign of the operation's type selects the instruction for the operations where it
matters:

| Operator | Signed type | Unsigned type |
|---|---|---|
| `/`, `//` | truncating signed division | unsigned division |
| `%` | remainder with the dividend's sign | unsigned remainder |
| `>>` | arithmetic shift (sign bit shifts in) | logical shift (zeros shift in) |
| `<`, `<=`, `>`, `>=` | signed comparison | unsigned comparison |

`+`, `-`, `*`, `<<`, `&`, `|`, `^`, `==` and `!=` produce the same bits for either sign.

Unary minus on an unsigned value follows C too: `-x` is `2^N - x` in the value's type.

## Comparisons

A comparison converts both operands with the same usual arithmetic conversions and then
compares in the resulting type. `c_uint64(10) > c_int64(-1)` is therefore an unsigned
comparison in which `-1` is the largest possible value, and the result is false. The
result of a comparison is `1` or `0`, as in C.

## A verifier gotcha: packet pointer fields

A few context fields are declared as 32-bit integers but are pointers as far as the
kernel verifier is concerned: `data`, `data_end` and `data_meta` on `xdp_md`, and `data`
and `data_end` on `__sk_buff`. Because they are `c_uint32`, arithmetic on them directly
is a 32-bit operation, exactly as in C, and the verifier rejects 32-bit arithmetic on a
pointer:

```
R0 32-bit pointer arithmetic prohibited
```

C programs cast these fields through `(void *)(long)` before using them for the same
reason. Until PythonBPF does this for you, copy the field into a local first, which is a
64-bit slot, or cast it with `c_void_p`:

```python
data = ctx.data # 64-bit local
end = ctx.data_end
if data + 34 < end: # 64-bit pointer arithmetic, accepted
...
```

```{note}
This is a known gap. The plan is to give these fields pointer rank automatically so that
no cast or copy is needed; this section will go away when that lands.
```

## Divergences from Python

Because the semantics are C's, some Python behaviour does not carry over:

* `/` is integer division; there is no floating-point result.
* `//` and `/` are the same operation, and both truncate toward zero: `-7 // 2` is `-3`,
where Python gives `-4`.
* `%` takes the sign of the dividend: `-7 % 2` is `-1`, where Python gives `1`.
* Integers have a fixed width and wrap on overflow; there is no arbitrary precision.
* Unsigned types exist, and mixing them with signed values follows C's conversions
rather than Python's mathematical integers.

The test programs under `tests/passing_tests/signedness/` show each rule with its
expected value, and `tests/c-form/signedness.bpf.c` is the equivalent C program.
29 changes: 22 additions & 7 deletions pythonbpf/allocation_pass.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,8 @@
from pythonbpf.helper import HelperHandlerRegistry
from pythonbpf.vmlinux_parser.dependency_node import Field
from .expr import VmlinuxHandlerRegistry
from pythonbpf.type_deducer import ctypes_to_ir
from pythonbpf.type_deducer import ctypes_to_ir, IntTy, signedness
from pythonbpf.expr.type_inference import infer_int_type
from pythonbpf.maps import BPFMapType

logger = logging.getLogger(__name__)
Expand Down Expand Up @@ -74,7 +75,9 @@ def handle_assign_allocation(compilation_context, builder, stmt, local_sym_tab):
elif isinstance(rval, ast.Constant):
_allocate_for_constant(builder, var_name, rval, local_sym_tab)
elif isinstance(rval, ast.BinOp):
_allocate_for_binop(builder, var_name, local_sym_tab)
_allocate_for_binop(
builder, var_name, rval, local_sym_tab, compilation_context
)
elif isinstance(rval, ast.Name):
# Variable-to-variable assignment (b = a)
_allocate_for_name(
Expand Down Expand Up @@ -116,7 +119,11 @@ def _allocate_for_call(builder, var_name, rval, local_sym_tab, compilation_conte

# Helper functions
elif HelperHandlerRegistry.has_handler(call_type):
ir_type = ir.IntType(64) # Assume i64 return type
# Undeclared locals are 64-bit; the sign comes from the helper.
ret = HelperHandlerRegistry.get_return_type(call_type)
ir_type = IntTy(
64, signedness(ret) if isinstance(ret, ir.IntType) else True
)
var = builder.alloca(ir_type, name=var_name)
var.align = 8
local_sym_tab[var_name] = LocalSymbol(var, ir_type)
Expand Down Expand Up @@ -256,7 +263,7 @@ def _allocate_for_constant(builder, var_name, rval, local_sym_tab):
"""Allocate memory for variable assigned from a constant."""

if isinstance(rval.value, bool):
ir_type = ir.IntType(1)
ir_type = IntTy(1, False) # a bool widens to 0 or 1, never sign-extends
var = builder.alloca(ir_type, name=var_name)
var.align = 1
local_sym_tab[var_name] = LocalSymbol(var, ir_type)
Expand All @@ -282,9 +289,17 @@ def _allocate_for_constant(builder, var_name, rval, local_sym_tab):
)


def _allocate_for_binop(builder, var_name, local_sym_tab):
"""Allocate memory for variable assigned from a binary operation."""
ir_type = ir.IntType(64) # Assume i64 result
def _allocate_for_binop(builder, var_name, rval, local_sym_tab, compilation_context):
"""Allocate memory for variable assigned from a binary operation.

Undeclared locals are 64-bit; the sign is that of the expression's C type,
inferred statically. Falls back to signed when the expression involves
something the inference does not know.
"""
inferred = infer_int_type(rval, local_sym_tab, compilation_context)
if inferred is None:
logger.debug(f"Could not infer a type for {var_name}, assuming signed i64")
ir_type = IntTy(64, signedness(inferred) if inferred is not None else True)
var = builder.alloca(ir_type, name=var_name)
var.align = 8
local_sym_tab[var_name] = LocalSymbol(var, ir_type)
Expand Down
22 changes: 8 additions & 14 deletions pythonbpf/assign_pass.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
from inspect import isclass

from llvmlite import ir
from pythonbpf.expr import eval_expr
from pythonbpf.expr import eval_expr, convert
from pythonbpf.helper import emit_probe_read_kernel_str_call
from pythonbpf.type_deducer import ctypes_to_ir
from pythonbpf.vmlinux_parser.dependency_node import Field
Expand Down Expand Up @@ -57,10 +57,7 @@ def handle_struct_field_assignment(
# Same implicit widening/truncation as assignment to a local: expressions
# evaluate in i64, but a field may be narrower.
if isinstance(val_type, ir.IntType) and isinstance(field_type, ir.IntType):
if val_type.width < field_type.width:
val = builder.sext(val, field_type)
elif val_type.width > field_type.width:
val = builder.trunc(val, field_type)
val = convert(builder, val, val_type, field_type)

# Regular assignment
builder.store(val, field_ptr)
Expand Down Expand Up @@ -154,7 +151,12 @@ def handle_variable_assignment(
f"Evaluated value for {var_name}: {val} of type {val_type}, expected {var_type}"
)

if val_type != var_type:
if isinstance(val_type, ir.IntType) and isinstance(var_type, ir.IntType):
# The descriptor may be narrower than the constant carrying the value
# (a literal is a 64-bit constant typed as C int), so never decide
# from descriptor equality: convert is a no-op when widths match.
val = convert(builder, val, val_type, var_type)
elif val_type != var_type:
# Handle vmlinux struct pointers - they're represented as Python classes but are i64 pointers
if isclass(val_type) and (val_type.__module__ == "vmlinux"):
logger.info("Handling vmlinux struct pointer assignment")
Expand Down Expand Up @@ -219,14 +221,6 @@ def handle_variable_assignment(
f"Failed to assign ctype struct field to {var_name}: {val_type} != {var_type}"
)
return False
elif isinstance(val_type, ir.IntType) and isinstance(var_type, ir.IntType):
# Allow implicit int widening
if val_type.width < var_type.width:
val = builder.sext(val, var_type)
logger.info(f"Implicitly widened int for variable {var_name}")
elif val_type.width > var_type.width:
val = builder.trunc(val, var_type)
logger.info(f"Implicitly truncated int for variable {var_name}")
elif isinstance(val_type, ir.IntType) and isinstance(var_type, ir.PointerType):
# NOTE: This is assignment to a PTR_TO_MAP_VALUE_OR_NULL
logger.info(
Expand Down
17 changes: 14 additions & 3 deletions pythonbpf/expr/__init__.py
Original file line number Diff line number Diff line change
@@ -1,5 +1,12 @@
from .expr_pass import eval_expr, handle_expr, get_operand_value
from .type_normalization import convert_to_bool, get_base_type_and_depth
from .expr_pass import eval_expr, handle_expr, get_typed_operand
from .type_normalization import (
convert_to_bool,
get_base_type_and_depth,
convert,
canonicalise,
to_promoted,
)
from .operators import usual_arithmetic_conversions
from .ir_ops import deref_to_depth, access_struct_field
from .operators import apply_binop
from .call_registry import CallHandlerRegistry
Expand All @@ -9,11 +16,15 @@
"eval_expr",
"handle_expr",
"convert_to_bool",
"convert",
"canonicalise",
"to_promoted",
"get_typed_operand",
"usual_arithmetic_conversions",
"get_base_type_and_depth",
"deref_to_depth",
"apply_binop",
"access_struct_field",
"get_operand_value",
"CallHandlerRegistry",
"VmlinuxHandlerRegistry",
]
Loading
Loading