Skip to content

Commit dc628ee

Browse files
authored
Merge pull request #101 from pythonbpf/feat/integer-signedness
integer signedness made consistent.
2 parents 3179549 + f727bfc commit dc628ee

35 files changed

Lines changed: 1457 additions & 218 deletions

‎docs/index.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -69,6 +69,7 @@ user-guide/maps
6969
user-guide/structs
7070
user-guide/compilation
7171
user-guide/helpers
72+
user-guide/integers
7273
```
7374

7475
```{toctree}

‎docs/user-guide/index.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -44,6 +44,9 @@ PythonBPF uses Python's `ctypes` module for type definitions:
4444
* `c_void_p` - Void pointers
4545
* `str(N)` - Fixed-length strings (e.g., `str(16)` for 16-byte string)
4646

47+
Integers follow C's rules for width, sign, conversion and arithmetic; see
48+
{doc}`integers` for the details and the places where this differs from Python.
49+
4750
## Example Structure
4851

4952
A typical PythonBPF program follows this structure:

‎docs/user-guide/integers.md‎

Lines changed: 153 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,153 @@
1+
# Integer Semantics
2+
3+
PythonBPF programs are Python syntax, but the integers in them behave as C integers: the
4+
program runs in the kernel as BPF bytecode, where every value is a fixed-width machine
5+
word. This page describes the rules the compiler applies. They are C's rules, applied to
6+
the `ctypes` types you declare, so a program's arithmetic matches what the equivalent C
7+
program compiled with clang would compute.
8+
9+
```{note}
10+
This is one of the few places where PythonBPF deliberately differs from Python. Python
11+
integers have arbitrary precision and no unsigned types; BPF has neither. The
12+
[divergences from Python](#divergences-from-python) are listed at the end of this page.
13+
```
14+
15+
## Types
16+
17+
An integer's type is the `ctypes` type it was declared with, and the type carries both
18+
a width and a sign:
19+
20+
| Signed | Unsigned | Width |
21+
|---|---|---|
22+
| `c_int8` | `c_uint8` | 8 |
23+
| `c_int16` | `c_uint16` | 16 |
24+
| `c_int32` | `c_uint32` | 32 |
25+
| `c_int64` | `c_uint64` | 64 |
26+
27+
Every declaration site uses these types: local variables initialised with a constructor
28+
call, `@bpfglobal` variables, `@struct` fields, map keys and values, and fields read from
29+
`vmlinux` structures. Helper functions return the type of the kernel's signature, so
30+
`pid()` and `ktime()` are unsigned while `probe_read`-style helpers return a signed
31+
`long`.
32+
33+
```python
34+
count = c_uint32(0) # a 32-bit unsigned local
35+
delta = c_int64(-1) # a 64-bit signed local
36+
```
37+
38+
A local assigned without a constructor takes its type from the expression:
39+
40+
```python
41+
now = ktime() # c_uint64, the helper's return type
42+
total = count + 1 # the type of the addition (see below), held in a 64-bit slot
43+
```
44+
45+
Undeclared locals are always 64 bits wide; the inferred type only decides their sign.
46+
Declare the local with a constructor when a narrower width matters.
47+
48+
### Literals
49+
50+
A literal has the type a C compiler gives it: `int` (32-bit signed) if the value fits,
51+
`long long` (64-bit signed) otherwise. This matters for mixed arithmetic: in
52+
`count / -2` with `count` a `c_uint32`, the literal `-2` is a 32-bit `int`, so the
53+
division happens in `c_uint32` exactly as it would in C.
54+
55+
## Assignment and conversion
56+
57+
Assigning a value to a variable of a different integer type converts it, and the
58+
variable's declared type is what the stored value *is* afterwards:
59+
60+
* **Widening preserves the value.** The conversion looks at the *source*'s sign: an
61+
unsigned source is zero-extended, a signed source is sign-extended. So a `c_uint32`
62+
holding `0xFFFFFFFF` stored into a `c_int64` gives `4294967295`, and a `c_int32`
63+
holding `-1` stored into a `c_uint64` gives `0xFFFFFFFFFFFFFFFF`. This is what C and
64+
`ctypes` both do.
65+
* **Narrowing truncates.** Only the low bits survive.
66+
* **Same width reinterprets.** A `c_uint32` `0xFFFFFFFF` stored into a `c_int32` reads
67+
as `-1`.
68+
69+
The same rules apply to explicit conversions written as constructor calls
70+
(`c_int64(count)`), to struct field stores and to `return`.
71+
72+
## Arithmetic
73+
74+
Each binary operation is typed on its own, from its two operands, following C's usual
75+
arithmetic conversions:
76+
77+
1. Operands narrower than 32 bits are promoted to `c_int32`.
78+
2. If both operands have the same sign, the result has the wider width and that sign.
79+
3. If the signs differ, the unsigned type wins when it is at least as wide as the signed
80+
one; otherwise the signed type wins.
81+
82+
The operation is then performed in that type, and its result has that type. The variable
83+
receiving the result plays no part until the final store. Two consequences worth knowing:
84+
85+
* **Intermediate results wrap at their own width.** `c_uint32(0x80000000) * c_uint32(2)`
86+
is a `c_uint32` multiplication, so it wraps to `0` before being stored, even if the
87+
destination is a `c_uint64`. Widen an operand first if you want a 64-bit product.
88+
* **Mixed signs go unsigned.** `c_uint32(10) / c_int32(-2)` is an unsigned division by
89+
`0xFFFFFFFE`, giving `0`, not `-5`.
90+
91+
The sign of the operation's type selects the instruction for the operations where it
92+
matters:
93+
94+
| Operator | Signed type | Unsigned type |
95+
|---|---|---|
96+
| `/`, `//` | truncating signed division | unsigned division |
97+
| `%` | remainder with the dividend's sign | unsigned remainder |
98+
| `>>` | arithmetic shift (sign bit shifts in) | logical shift (zeros shift in) |
99+
| `<`, `<=`, `>`, `>=` | signed comparison | unsigned comparison |
100+
101+
`+`, `-`, `*`, `<<`, `&`, `|`, `^`, `==` and `!=` produce the same bits for either sign.
102+
103+
Unary minus on an unsigned value follows C too: `-x` is `2^N - x` in the value's type.
104+
105+
## Comparisons
106+
107+
A comparison converts both operands with the same usual arithmetic conversions and then
108+
compares in the resulting type. `c_uint64(10) > c_int64(-1)` is therefore an unsigned
109+
comparison in which `-1` is the largest possible value, and the result is false. The
110+
result of a comparison is `1` or `0`, as in C.
111+
112+
## A verifier gotcha: packet pointer fields
113+
114+
A few context fields are declared as 32-bit integers but are pointers as far as the
115+
kernel verifier is concerned: `data`, `data_end` and `data_meta` on `xdp_md`, and `data`
116+
and `data_end` on `__sk_buff`. Because they are `c_uint32`, arithmetic on them directly
117+
is a 32-bit operation, exactly as in C, and the verifier rejects 32-bit arithmetic on a
118+
pointer:
119+
120+
```
121+
R0 32-bit pointer arithmetic prohibited
122+
```
123+
124+
C programs cast these fields through `(void *)(long)` before using them for the same
125+
reason. Until PythonBPF does this for you, copy the field into a local first, which is a
126+
64-bit slot, or cast it with `c_void_p`:
127+
128+
```python
129+
data = ctx.data # 64-bit local
130+
end = ctx.data_end
131+
if data + 34 < end: # 64-bit pointer arithmetic, accepted
132+
...
133+
```
134+
135+
```{note}
136+
This is a known gap. The plan is to give these fields pointer rank automatically so that
137+
no cast or copy is needed; this section will go away when that lands.
138+
```
139+
140+
## Divergences from Python
141+
142+
Because the semantics are C's, some Python behaviour does not carry over:
143+
144+
* `/` is integer division; there is no floating-point result.
145+
* `//` and `/` are the same operation, and both truncate toward zero: `-7 // 2` is `-3`,
146+
where Python gives `-4`.
147+
* `%` takes the sign of the dividend: `-7 % 2` is `-1`, where Python gives `1`.
148+
* Integers have a fixed width and wrap on overflow; there is no arbitrary precision.
149+
* Unsigned types exist, and mixing them with signed values follows C's conversions
150+
rather than Python's mathematical integers.
151+
152+
The test programs under `tests/passing_tests/signedness/` show each rule with its
153+
expected value, and `tests/c-form/signedness.bpf.c` is the equivalent C program.

‎pythonbpf/allocation_pass.py‎

Lines changed: 22 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,8 @@
66
from pythonbpf.helper import HelperHandlerRegistry
77
from pythonbpf.vmlinux_parser.dependency_node import Field
88
from .expr import VmlinuxHandlerRegistry
9-
from pythonbpf.type_deducer import ctypes_to_ir
9+
from pythonbpf.type_deducer import ctypes_to_ir, IntTy, signedness
10+
from pythonbpf.expr.type_inference import infer_int_type
1011
from pythonbpf.maps import BPFMapType
1112

1213
logger = logging.getLogger(__name__)
@@ -74,7 +75,9 @@ def handle_assign_allocation(compilation_context, builder, stmt, local_sym_tab):
7475
elif isinstance(rval, ast.Constant):
7576
_allocate_for_constant(builder, var_name, rval, local_sym_tab)
7677
elif isinstance(rval, ast.BinOp):
77-
_allocate_for_binop(builder, var_name, local_sym_tab)
78+
_allocate_for_binop(
79+
builder, var_name, rval, local_sym_tab, compilation_context
80+
)
7881
elif isinstance(rval, ast.Name):
7982
# Variable-to-variable assignment (b = a)
8083
_allocate_for_name(
@@ -116,7 +119,11 @@ def _allocate_for_call(builder, var_name, rval, local_sym_tab, compilation_conte
116119

117120
# Helper functions
118121
elif HelperHandlerRegistry.has_handler(call_type):
119-
ir_type = ir.IntType(64) # Assume i64 return type
122+
# Undeclared locals are 64-bit; the sign comes from the helper.
123+
ret = HelperHandlerRegistry.get_return_type(call_type)
124+
ir_type = IntTy(
125+
64, signedness(ret) if isinstance(ret, ir.IntType) else True
126+
)
120127
var = builder.alloca(ir_type, name=var_name)
121128
var.align = 8
122129
local_sym_tab[var_name] = LocalSymbol(var, ir_type)
@@ -256,7 +263,7 @@ def _allocate_for_constant(builder, var_name, rval, local_sym_tab):
256263
"""Allocate memory for variable assigned from a constant."""
257264

258265
if isinstance(rval.value, bool):
259-
ir_type = ir.IntType(1)
266+
ir_type = IntTy(1, False) # a bool widens to 0 or 1, never sign-extends
260267
var = builder.alloca(ir_type, name=var_name)
261268
var.align = 1
262269
local_sym_tab[var_name] = LocalSymbol(var, ir_type)
@@ -282,9 +289,17 @@ def _allocate_for_constant(builder, var_name, rval, local_sym_tab):
282289
)
283290

284291

285-
def _allocate_for_binop(builder, var_name, local_sym_tab):
286-
"""Allocate memory for variable assigned from a binary operation."""
287-
ir_type = ir.IntType(64) # Assume i64 result
292+
def _allocate_for_binop(builder, var_name, rval, local_sym_tab, compilation_context):
293+
"""Allocate memory for variable assigned from a binary operation.
294+
295+
Undeclared locals are 64-bit; the sign is that of the expression's C type,
296+
inferred statically. Falls back to signed when the expression involves
297+
something the inference does not know.
298+
"""
299+
inferred = infer_int_type(rval, local_sym_tab, compilation_context)
300+
if inferred is None:
301+
logger.debug(f"Could not infer a type for {var_name}, assuming signed i64")
302+
ir_type = IntTy(64, signedness(inferred) if inferred is not None else True)
288303
var = builder.alloca(ir_type, name=var_name)
289304
var.align = 8
290305
local_sym_tab[var_name] = LocalSymbol(var, ir_type)

‎pythonbpf/assign_pass.py‎

Lines changed: 8 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@
33
from inspect import isclass
44

55
from llvmlite import ir
6-
from pythonbpf.expr import eval_expr
6+
from pythonbpf.expr import eval_expr, convert
77
from pythonbpf.helper import emit_probe_read_kernel_str_call
88
from pythonbpf.type_deducer import ctypes_to_ir
99
from pythonbpf.vmlinux_parser.dependency_node import Field
@@ -57,10 +57,7 @@ def handle_struct_field_assignment(
5757
# Same implicit widening/truncation as assignment to a local: expressions
5858
# evaluate in i64, but a field may be narrower.
5959
if isinstance(val_type, ir.IntType) and isinstance(field_type, ir.IntType):
60-
if val_type.width < field_type.width:
61-
val = builder.sext(val, field_type)
62-
elif val_type.width > field_type.width:
63-
val = builder.trunc(val, field_type)
60+
val = convert(builder, val, val_type, field_type)
6461

6562
# Regular assignment
6663
builder.store(val, field_ptr)
@@ -154,7 +151,12 @@ def handle_variable_assignment(
154151
f"Evaluated value for {var_name}: {val} of type {val_type}, expected {var_type}"
155152
)
156153

157-
if val_type != var_type:
154+
if isinstance(val_type, ir.IntType) and isinstance(var_type, ir.IntType):
155+
# The descriptor may be narrower than the constant carrying the value
156+
# (a literal is a 64-bit constant typed as C int), so never decide
157+
# from descriptor equality: convert is a no-op when widths match.
158+
val = convert(builder, val, val_type, var_type)
159+
elif val_type != var_type:
158160
# Handle vmlinux struct pointers - they're represented as Python classes but are i64 pointers
159161
if isclass(val_type) and (val_type.__module__ == "vmlinux"):
160162
logger.info("Handling vmlinux struct pointer assignment")
@@ -219,14 +221,6 @@ def handle_variable_assignment(
219221
f"Failed to assign ctype struct field to {var_name}: {val_type} != {var_type}"
220222
)
221223
return False
222-
elif isinstance(val_type, ir.IntType) and isinstance(var_type, ir.IntType):
223-
# Allow implicit int widening
224-
if val_type.width < var_type.width:
225-
val = builder.sext(val, var_type)
226-
logger.info(f"Implicitly widened int for variable {var_name}")
227-
elif val_type.width > var_type.width:
228-
val = builder.trunc(val, var_type)
229-
logger.info(f"Implicitly truncated int for variable {var_name}")
230224
elif isinstance(val_type, ir.IntType) and isinstance(var_type, ir.PointerType):
231225
# NOTE: This is assignment to a PTR_TO_MAP_VALUE_OR_NULL
232226
logger.info(

‎pythonbpf/expr/__init__.py‎

Lines changed: 14 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,12 @@
1-
from .expr_pass import eval_expr, handle_expr, get_operand_value
2-
from .type_normalization import convert_to_bool, get_base_type_and_depth
1+
from .expr_pass import eval_expr, handle_expr, get_typed_operand
2+
from .type_normalization import (
3+
convert_to_bool,
4+
get_base_type_and_depth,
5+
convert,
6+
canonicalise,
7+
to_promoted,
8+
)
9+
from .operators import usual_arithmetic_conversions
310
from .ir_ops import deref_to_depth, access_struct_field
411
from .operators import apply_binop
512
from .call_registry import CallHandlerRegistry
@@ -9,11 +16,15 @@
916
"eval_expr",
1017
"handle_expr",
1118
"convert_to_bool",
19+
"convert",
20+
"canonicalise",
21+
"to_promoted",
22+
"get_typed_operand",
23+
"usual_arithmetic_conversions",
1224
"get_base_type_and_depth",
1325
"deref_to_depth",
1426
"apply_binop",
1527
"access_struct_field",
16-
"get_operand_value",
1728
"CallHandlerRegistry",
1829
"VmlinuxHandlerRegistry",
1930
]

0 commit comments

Comments
 (0)