Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/agents/algo-implementer.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,5 +28,5 @@ from its `roadmap/<domain>/<Name>/` spec into the codebase, following both exact
- **No comments** (except license/`NOLINT`). Allman braces, brace-init.
- **Tests**: `TEST_F` on `float`, `StrictMock` only, anonymous-namespace fixture; implement EXACTLY
the spec's cases — no redundant or extra tests.
- **Embedded**: scoped `#pragma GCC push_options`/`optimize`/`pop_options` + `OPTIMIZE_FOR_SPEED` on hot paths.
- **Embedded**: no `#pragma GCC optimize`/`optimize` attribute; `OPTIMIZE_FOR_SPEED` (forced inlining) on hot paths.
- **Terse**: no preamble/postamble, no plan restatement, no narration; don't re-read files; batch reads.
2 changes: 1 addition & 1 deletion .claude/agents/executor.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: executor
description: Implement code changes in numerical-toolbox — float-only templates, no heap, embedded pragmas, TEST_F on float, CMake wiring, docs. Needs a clear task or plan.
description: Implement code changes in numerical-toolbox — float-only templates, no heap, forced inlining on hot paths, TEST_F on float, CMake wiring, docs. Needs a clear task or plan.
model: claude-sonnet-4-6
tools: [Read, Write, Edit, Bash, TodoWrite]
---
Expand Down
2 changes: 1 addition & 1 deletion .claude/agents/planner.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Canonical rules: `AGENTS.md`. Produce plans only — no code edits.
- [ ] No heap; no recursion; tests too
- [ ] `template<typename T>` + `static_assert(std::is_floating_point_v<T>)`; `float` only
- [ ] `TEST_F` on `float` — no `TYPED_TEST`; `StrictMock` only; never plain `TEST()`
- [ ] scoped `#pragma GCC push_options`/`optimize`/`pop_options` + `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] no `#pragma GCC optimize`/`optimize` attribute; `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] `doc/` update planned

**Terse**: no preamble/postamble, no plan restatement; don't re-read files; batch reads.
4 changes: 2 additions & 2 deletions .claude/agents/reviewer.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: reviewer
description: Review code changes against numerical-toolbox standards — no heap, float-only templates, embedded pragmas, TEST_F on float, SOLID, docs. Does NOT modify files.
description: Review code changes against numerical-toolbox standards — no heap, float-only templates, forced inlining on hot paths, TEST_F on float, SOLID, docs. Does NOT modify files.
model: claude-sonnet-4-6
tools: [Read, Bash]
---
Expand Down Expand Up @@ -35,7 +35,7 @@ End with totals + verdict: APPROVE / REQUEST CHANGES.
- [ ] `extern template` guarded by `#ifdef NUMERICAL_TOOLBOX_COVERAGE_BUILD`.

**Embedded optimizations (WARNING)**
- [ ] `#pragma GCC push_options` + `optimize("O3","fast-math")` after `#pragma once`, matching `pop_options` at end of algorithm headers (no TU leak; none in test `.cpp`).
- [ ] No `#pragma GCC optimize` and no `optimize` attribute anywhere (they stop GCC inlining across the boundary).
- [ ] `OPTIMIZE_FOR_SPEED` on `Filter/Compute/Update/Solve/Step`.

**Namespaces (WARNING)**
Expand Down
2 changes: 1 addition & 1 deletion .github/agents/algo-implementer.agent.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,5 +31,5 @@ Authoritative rules: `AGENTS.md`. Recipe: `roadmap/DEPLOYMENT.md`. Follow both e
- **No comments** (except license/`NOLINT`). Allman braces, brace-init.
- **Tests**: `TEST_F` on `float`, `StrictMock` only, anonymous-namespace fixture; implement EXACTLY
the spec's cases — no redundant or extra tests.
- **Embedded**: scoped `#pragma GCC push_options`/`optimize`/`pop_options` + `OPTIMIZE_FOR_SPEED` on hot paths.
- **Embedded**: no `#pragma GCC optimize`/`optimize` attribute; `OPTIMIZE_FOR_SPEED` (forced inlining) on hot paths.
- **Terse**: no preamble/postamble, no plan restatement, no narration; don't re-read files; batch reads.
2 changes: 1 addition & 1 deletion .github/agents/executor.agent.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
description: "Implement code changes in numerical-toolbox — float-only templates, no heap, embedded pragmas, TEST_F on float, CMake wiring, docs. Needs a clear task or plan."
description: "Implement code changes in numerical-toolbox — float-only templates, no heap, forced inlining on hot paths, TEST_F on float, CMake wiring, docs. Needs a clear task or plan."
tools: [read, edit, search, execute, todo]
model: "Claude Sonnet 4.6"
handoffs:
Expand Down
2 changes: 1 addition & 1 deletion .github/agents/planner.agent.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ Canonical rules: `AGENTS.md`. Produce plans only — no code edits.
- [ ] No heap; no recursion; tests too
- [ ] `template<typename T>` + `static_assert(std::is_floating_point_v<T>)`; `float` only
- [ ] `TEST_F` on `float` — no `TYPED_TEST`; `StrictMock` only; never plain `TEST()`
- [ ] scoped `#pragma GCC push_options`/`optimize`/`pop_options` + `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] no `#pragma GCC optimize`/`optimize` attribute; `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] `doc/` update planned

**Terse**: no preamble/postamble, no plan restatement; don't re-read files; batch reads.
4 changes: 2 additions & 2 deletions .github/agents/reviewer.agent.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
description: "Review code changes against numerical-toolbox standards: no heap, float-only templates, embedded pragmas, TEST_F on float, SOLID, docs. Does NOT modify files."
description: "Review code changes against numerical-toolbox standards: no heap, float-only templates, forced inlining on hot paths, TEST_F on float, SOLID, docs. Does NOT modify files."
tools: [read, search]
model: "claude-sonnet-4-6"
handoffs:
Expand Down Expand Up @@ -41,7 +41,7 @@ End with totals + verdict: APPROVE / REQUEST CHANGES.
- [ ] `extern template` guarded by `#ifdef NUMERICAL_TOOLBOX_COVERAGE_BUILD`.

**Embedded optimizations (WARNING)**
- [ ] `#pragma GCC push_options` + `optimize("O3","fast-math")` after `#pragma once`, matching `pop_options` at end of algorithm headers (no TU leak; none in test `.cpp`).
- [ ] No `#pragma GCC optimize` and no `optimize` attribute anywhere (they stop GCC inlining across the boundary).
- [ ] `OPTIMIZE_FOR_SPEED` on `Filter/Compute/Update/Solve/Step`.

**Namespaces (WARNING)**
Expand Down
2 changes: 1 addition & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Code-file specifics: `.github/instructions/` (applyTo-scoped). Deployment recipe
Essentials (full detail in AGENTS.md):
- **No heap** — bounded containers / `std::array` / `std::optional`; no recursion; tests too.
- **Float-only** — generic `template<typename T>` + `static_assert(std::is_floating_point_v<T>)`; instantiate/test `float`; no Q15/Q31.
- **Embedded** — scoped `#pragma GCC push_options` / `optimize("O3","fast-math")` … `pop_options` (never leak to the TU) + `OPTIMIZE_FOR_SPEED` on hot paths.
- **Embedded** — no `#pragma GCC optimize` / `optimize` attribute (they block inlining); the consumer sets optimization flags per TU; `OPTIMIZE_FOR_SPEED` (forced inlining) on hot paths.
- **No comments** (except license/NOLINT). Allman braces, brace-init, PascalCase/camelCase.
- **Tests** — `TEST_F` on `float`, `StrictMock` only, never plain `TEST()`, no redundant cases.
- **No exceptions** — `std::optional`/status enums; interfaces `virtual ~I() = default`.
Expand Down
23 changes: 4 additions & 19 deletions .github/instructions/numerical-cpp.instructions.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
description: "Numerical C++ rules (float-only): no heap, bounded containers, generic template<typename T> instantiated for float, embedded pragmas, Allman/brace-init, SOLID, const-correct. Canonical: AGENTS.md."
description: "Numerical C++ rules (float-only): no heap, bounded containers, generic template<typename T> instantiated for float, forced inlining on hot paths, Allman/brace-init, SOLID, const-correct. Canonical: AGENTS.md."
applyTo: "**/*.{hpp,cpp,h}"
---

Expand Down Expand Up @@ -41,26 +41,11 @@ To replace a function with a platform-specific implementation, define the corres

## Embedded Optimizations

Every algorithm header MUST bracket its body with a scoped pragma, so the options never leak into the including translation unit:
Never use `#pragma GCC optimize` or the `optimize` attribute. GCC does not inline a callee whose optimization options differ from its caller's, so options set per header or per function turn every small helper on a hot path into an out-of-line call. Optimization level and floating-point flags are the consumer's, applied to whole translation units (embedded: `-O2`/`-O3` with `-ffast-math -fno-finite-math-only`).

```cpp
#pragma once
Apply `OPTIMIZE_FOR_SPEED` (from `numerical/math/CompilerOptimizations.hpp`) on hot-path methods: `Compute()`, `Filter()`, `Calculate()`, `Solve()`, `Update()`, `Step()`. It forces inlining when `NumericalToolbox_ENABLE_OPTIMIZATIONS` is defined and expands to nothing otherwise.

#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC push_options
#pragma GCC optimize("O3", "fast-math")
#endif

// ... includes and header body ...

#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC pop_options
#endif
```

The `pop_options` block is the last thing in the header. Never use the pragma in test `.cpp` files.

Apply `OPTIMIZE_FOR_SPEED` (from `numerical/math/CompilerOptimizations.hpp`) on hot-path methods: `Compute()`, `Filter()`, `Calculate()`, `Solve()`, `Update()`, `Step()`.
Use `math::IsFinite` for finiteness checks: a consumer may still build with `-ffinite-math-only`.

## Naming

Expand Down
35 changes: 12 additions & 23 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,29 +22,18 @@ resource-constrained embedded systems. Real-time, deterministic, no heap.
`infra::BoundedDeque<T>::WithMaxSize<N>`, `infra::BoundedList<T>::WithMaxSize<N>`,
`std::array<T,N>`, `std::optional<T>`. Stack/static only. No recursion. **Tests too.**

## Embedded optimizations (algorithm headers)

```cpp
#pragma once
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC push_options
#pragma GCC optimize("O3", "fast-math")
#endif
#include "numerical/math/CompilerOptimizations.hpp"

// ... header body ...

#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC pop_options
#endif
```

The `pop_options` block is the **last thing in the header**: the pragma must never leak into the
including translation unit (it would silently apply fast-math, e.g. folded NaN checks and
reassociation, to unrelated code). Use `math::IsFinite` for finiteness checks. Test `.cpp` files
never use the pragma.

`OPTIMIZE_FOR_SPEED` on hot paths (`Filter/Compute/Update/Solve/Step`). Pure interfaces exempt.
## Embedded optimizations

- **No `#pragma GCC optimize` and no `optimize` attribute**, in headers or sources. GCC does not
inline a callee whose optimization options differ from its caller's, so options set per header or
per function turn every small helper on a hot path (element access, accessors, `math::` wrappers)
into an out-of-line call.
- The optimization level and floating-point model belong to the consumer and apply to whole
translation units. For embedded targets: `-O2` or `-O3` with `-ffast-math -fno-finite-math-only`.
- `OPTIMIZE_FOR_SPEED` on hot paths (`Filter/Compute/Update/Solve/Step`). Pure interfaces exempt.
With `NumericalToolbox_ENABLE_OPTIMIZATIONS` it forces inlining (`always_inline`, `hot`, `inline`);
without it, it expands to nothing.
- Use `math::IsFinite` for finiteness checks: a consumer may still build with `-ffinite-math-only`.

## Style

Expand Down
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Code-file specifics: `.github/instructions/`. Deployment recipe: `roadmap/DEPLOY
Essentials (full detail in AGENTS.md):
- **No heap** — bounded containers / `std::array` / `std::optional`; no recursion; tests too.
- **Float-only** — generic `template<typename T>` + `static_assert(std::is_floating_point_v<T>)`; instantiate/test `float`; no Q15/Q31.
- **Embedded** — scoped `#pragma GCC push_options` / `optimize("O3","fast-math")` … `pop_options` (never leak to the TU) + `OPTIMIZE_FOR_SPEED` on hot paths.
- **Embedded** — no `#pragma GCC optimize` / `optimize` attribute (they block inlining); the consumer sets optimization flags per TU; `OPTIMIZE_FOR_SPEED` (forced inlining) on hot paths.
- **No comments** (except license/NOLINT). Allman braces, brace-init, PascalCase/camelCase.
- **Tests** — `TEST_F` on `float`, `StrictMock` only, never plain `TEST()`, no redundant cases.
- **No exceptions** — `std::optional`/status enums; interfaces `virtual ~I() = default`.
Expand Down
4 changes: 2 additions & 2 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
Prioritized backlog of reusable numerical components for generic embedded applications,
ordered **easiest → hardest to implement** within the constraints of this library
(templated on `float` / `math::Q15` / `math::Q31`, no heap, bounded containers,
`#pragma GCC optimize` + `OPTIMIZE_FOR_SPEED` hot paths, typed tests, `doc/` page).
`OPTIMIZE_FOR_SPEED` hot paths, typed tests, `doc/` page).

Difficulty legend:

Expand Down Expand Up @@ -52,7 +52,7 @@ on bounded `math::Vector`/`math::Matrix` inputs; tests are `TEST_F` on `float`.
Every new component should follow the established repository conventions:

- [ ] Header-only template supporting `float` / `math::Q15` / `math::Q31` (or *float-first* where noted)
- [ ] Scoped `#pragma GCC push_options` / `optimize("O3", "fast-math")` after `#pragma once`, `pop_options` at end of file; `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] No `#pragma GCC optimize` / `optimize` attribute; `OPTIMIZE_FOR_SPEED` on hot paths
- [ ] No heap, no recursion, bounded containers (`infra::BoundedVector`, `std::array`)
- [ ] `static_assert` on supported types and dimensions
- [ ] Typed tests (`TYPED_TEST`) for multi-type components; `TEST_F` for single-type; `StrictMock` only
Expand Down
4 changes: 4 additions & 0 deletions cmake/QemuHelpers.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -11,4 +11,8 @@ function(numerical_link_qemu_runtime target)
hal.qemu.cortex
gmock_main
)
# newlib-nano's printf has no float conversions unless _printf_float is linked. Without them the
# vsnprintf behind std::ostream << float returns a bogus length, libstdc++ allocas that many bytes,
# and the first failure message that prints a float wraps the stack pointer and locks up the core.
target_link_options(${target} PRIVATE "LINKER:--undefined=_printf_float")
endfunction()
2 changes: 1 addition & 1 deletion doc/analysis/DiscreteWaveletTransform.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ $$\hat{x} = [1.0,\; 2.0,\; 3.0,\; 4.0] \checkmark$$
- **Filter length vs. signal length.** At each level the signal halves; once it equals $P$ the
periodic convolution wraps completely. Stop decomposition before the signal shorter than the
filter length to avoid artefacts.
- **Fast-math reordering.** With `#pragma GCC optimize("fast-math")` floating-point associativity
- **Fast-math reordering.** With `-ffast-math` floating-point associativity
relaxes; reconstruction residuals may reach $10^{-5}$ rather than $10^{-7}$ for 32-bit floats.

## Variants & Generalizations
Expand Down
4 changes: 2 additions & 2 deletions doc/filters/passive/BiquadCascade.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,8 +94,8 @@ Input $[1, 0, 0, \ldots]$ → output $[1, 0, 0, \ldots]$ — identity passthroug
$|z| = 1$. Finite-precision rounding can move a pole just outside, causing instability.
Use $Q \leq 30$ in single precision; double precision or lattice realizations for higher $Q$.
- **Denormal floats**: small state values approaching the denormal range stall the FPU pipeline
on many embedded cores. Enabling flush-to-zero (FTZ) or the fast-math pragma prevents this
at the cost of negligible numerical error.
on many embedded cores. Enabling flush-to-zero (FTZ) prevents this at the cost of negligible
numerical error.
- **DC gain normalization**: the RBJ low-pass has unity DC gain by construction. Gain-staging
between sections is not required; each section's output is well-scaled relative to its input.

Expand Down
2 changes: 1 addition & 1 deletion doc/math/MatrixNorms.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ Vector $\mathbf{v} = [3,\, 4]^\top$: $\|\mathbf{v}\|_2 = 5$, and $\hat{\mathbf{v

**Zero vector normalisation** — dividing by $\|\mathbf{v}\|_2 = 0$ is undefined. The implementation returns an empty optional for near-zero norms.

**Fast-math semantics** — `#pragma GCC optimize("fast-math")` may reorder floating-point operations. The norms are sums of non-negative values, so reordering does not change the sign of the result, but catastrophic cancellation can still occur for near-zero off-diagonal entries.
**Fast-math semantics** — `-ffast-math` may reorder floating-point operations. The norms are sums of non-negative values, so reordering does not change the sign of the result, but catastrophic cancellation can still occur for near-zero off-diagonal entries.

**Non-square matrices** — FrobeniusNorm, OneNorm, and InfinityNorm apply to any $m \times n$ matrix.

Expand Down
71 changes: 32 additions & 39 deletions doc/performance-optimization/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,19 +44,23 @@ set(CMAKE_EXE_LINKER_FLAGS "${CMAKE_EXE_LINKER_FLAGS} -Wl,--gc-sections")

### Fast-Math Considerations

```cpp
// Enables aggressive floating-point optimizations
// WARNING: May change numerical behavior slightly
#pragma GCC optimize("fast-math")
```cmake
# Aggressive floating-point optimizations for whole translation units, keeping NaN/Inf checks
add_compile_options(-ffast-math -fno-finite-math-only)
```

Effects of `-ffast-math`:
- Assumes no NaN or Infinity
- Allows reordering of operations
- No `errno` from math functions, so `sqrtf` becomes a single `vsqrt.f32`
- Allows reordering of operations and reciprocal multiplication instead of division
- Enables FMA (Fused Multiply-Add) instructions
- `-ffinite-math-only` (part of `-ffast-math`) assumes no NaN or Infinity and folds `std::isnan`/`std::isfinite`
- May break IEEE 754 compliance

**Use only when**: You control all inputs and don't need strict IEEE behavior.
**Use only when**: You control all inputs and don't need strict IEEE behavior. Keep
`-fno-finite-math-only` unless no code relies on NaN or Infinity checks.

Set these options for whole translation units, never per function: see
[Optimization Options Must Match Across Calls](#optimization-options-must-match-across-calls).

---

Expand Down Expand Up @@ -135,10 +139,17 @@ inline constexpr std::array<float, 512> sineLUT = []() {
#define HOT_FUNCTION __attribute__((hot))

// Combined macro for critical functions
#define OPTIMIZE_FOR_SPEED \
__attribute__((always_inline, hot, optimize("-O3"), optimize("-ffast-math"))) inline
#define OPTIMIZE_FOR_SPEED __attribute__((always_inline, hot)) inline
```

#### Optimization Options Must Match Across Calls

GCC does not inline a function into a caller whose optimization options differ from its own.
`#pragma GCC optimize` and `__attribute__((optimize(...)))` give the functions they cover options of
their own, so every small helper such a function calls — element access, accessors, unit wrappers —
stays an out-of-line call, and the functions themselves are no longer inlined into their callers.
Choose the optimization level and floating-point flags for whole translation units instead.

### 5. Prefer Fixed-Size Types

```cpp
Expand Down Expand Up @@ -204,32 +215,16 @@ set(CMAKE_CXX_FLAGS_DEBUG "-Og -g" CACHE STRING "Debug flags" FORCE)
- Register allocation
- Still debuggable (variable inspection works)

### Solution 2: Per-File Optimization Pragmas

```cpp
// Bracket the performance-critical code; never leave the options active for the rest of the TU
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC push_options
#pragma GCC optimize("O3", "fast-math")
#endif

// Performance-critical implementation...
### Solution 2: Per-Translation-Unit Options

#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC pop_options
#endif
```

### Solution 3: Per-Function Attributes

```cpp
__attribute__((optimize("-O3")))
void CriticalFunction() {
// This function is always optimized
}
```cmake
# Build the performance-critical sources optimized while the rest stays debuggable
set_source_files_properties(Controller.cpp PROPERTIES COMPILE_OPTIONS "$<$<CONFIG:Debug>:-O2>")
```

**Note**: Function-level attributes don't propagate to callees. Use file-level pragmas for better results.
Every function in a translation unit shares its options, so inlining is unaffected. Do not use
`#pragma GCC optimize` or `__attribute__((optimize(...)))` instead: see
[Optimization Options Must Match Across Calls](#optimization-options-must-match-across-calls).

---

Expand Down Expand Up @@ -430,13 +425,11 @@ struct GoodStruct {

## Quick Reference Card

### GCC Optimization Pragmas
```cpp
#pragma GCC optimize("O3") // Maximum speed
#pragma GCC optimize("Os") // Minimum size
#pragma GCC optimize("fast-math") // Aggressive FP
#pragma GCC push_options // Save current options
#pragma GCC pop_options // Restore options
### Optimization Flags (whole translation units)
```cmake
add_compile_options(-O3) # Maximum speed
add_compile_options(-Os) # Minimum size
add_compile_options(-ffast-math -fno-finite-math-only) # Aggressive FP, NaN/Inf checks kept
```

### Function Attributes
Expand Down
Loading
Loading