Skip to content

[Bug]: StdDev returns zero or NaN for float and double values offset far from zero #1196

Description

@dwcullop

Please note we can't commit to any timeline.

Describe the bug 🐞

The float and double StdDev selectors lose the variance entirely for values offset far from zero, returning 0, a large wrong value, or NaN instead of the correct sample standard deviation.

All selectors combine their accumulated moments as SumOfSquares - (SumOfItems * SumOfItems / Count). For the floating point selectors both terms are large and nearly equal, while their difference is small, so the subtraction is catastrophically cancelling. Once the true difference drops below the spacing of representable values at that magnitude it rounds away completely. When rounding pushes the difference below zero, the private Sqrt helper receives a negative argument and the result is NaN.

This is PRE-EXISTING behavior. It was neither introduced nor worsened by PR #1191. That PR corrected divisor placement and the integer path; the numerator subexpression is character for character identical before and after it, which was checked directly against 36ef6dba^.

The defect is not uniform across the five overloads:

Selector Behavior
int, long Exact as of PR #1191, which cross multiplies so the cancelling subtraction stays in integer arithmetic. Conditional on the long accumulators not overflowing; the final division and square root still round.
float Fails by cancellation, at ordinary magnitudes.
double Fails by cancellation, at larger magnitudes.
decimal Fails by a DIFFERENT mechanism, OverflowException, not cancellation.

Step to reproduce

Evidence below is from standalone C# arithmetic snippets that reproduce the selector's accumulate then combine sequence. They are NOT operator runs and NOT unit tests. No DynamicData pipeline or test fixture was executed for these findings.

double, three values one apart at an offset of 1e8. True sample deviation is 1.

double[] d = { 1e8, 1e8 + 1, 1e8 + 2 };
double sum = 0, sumSq = 0;
foreach (var x in d) { sum += x; sumSq += x * x; }
double numerator = sumSq - (sum * sum) / 3.0;
// sum         = 300000003         (exact)
// sumSq       = 30000000600000004
// numerator   = 0                 (exactly zero)
// Math.Sqrt(numerator / 2.0) = 0  (true answer: 1)

The inputs are distinct and exactly representable. The loss is in sumSq, which lands near 3.0e16 where consecutive doubles are 4 apart. The true difference between sumSq and sum * sum / Count is 2, below that spacing, so it rounds away and the variance collapses to exactly zero.

float, three values one apart at an offset of 1e5. True sample deviation is 1.

float[] f = { 100000f, 100001f, 100002f };
float sum = 0, sumSq = 0;
foreach (var x in f) { sum += x; sumSq += x * x; }
float numerator = sumSq - (sum * sum) / 3f;
// sum       = 300003    (exact)
// sumSq     = 30000600000   (true value is 30000600005)
// numerator = -2048
// Math.Sqrt(numerator / 2f) = NaN  (true answer: 1)

The inputs were checked explicitly for degeneracy, because an earlier candidate example was withdrawn for collapsing to a single float:

  • bit patterns 0x47C35000, 0x47C35080, 0x47C35100, all distinct
  • float spacing near 1e5 is 0.0078125, so integers one apart are exactly representable
  • each input compares exactly equal to its integer value
  • SumOfItems accumulates exactly to 300003

The loss is entirely in SumOfSquares, which lands near 3.0e10 where the float spacing is 2048. Five is lost from the true sum, the numerator evaluates to -2048, and the square root of a negative value is NaN.

Worth emphasising that float fails at an offset of 1e5, an ordinary magnitude, far below the 1e8 and above offsets needed to break double.

Frequency. NaN is reachable, and not only from contrived input. A seeded randomized sweep of the double combination (seed 20260922, 400000 trials, offsets 1e8 to 1e17, 3 to 7 items, spread 0 to 5) produced a negative numerator in 113790 trials, about 28 percent. This is common rather than exotic.

decimal, for completeness. Three values near 1e14 give SumOfItems * SumOfItems of about 9.0e28 against decimal.MaxValue of about 7.9228e28, and the expression throws OverflowException before any division. Separately, the private Sqrt helper in StdDevEx.cs throws OverflowException on a negative argument rather than returning NaN, so a decimal cancellation would surface as OnError rather than a quiet wrong value. Reordering to (SumOfItems / Count) * SumOfItems would raise that ceiling, but it carries its own rounding consequences and is not proposed here.

Reproduction repository

https://github.com/reactivemarbles/DynamicData/tree/3d2136021ae4ee71d0c352c983f6b0d7628ba759

The affected expression is in src/DynamicData/Aggregation/StdDevEx.cs.

Expected behavior

Equally spaced values should report their spacing regardless of how far the values are offset from zero. StdDev over [100000f, 100001f, 100002f] should return 1, not NaN, and StdDev over [1e8, 1e8 + 1, 1e8 + 2] should return 1, not 0. The result should never be NaN for finite input.

Screenshots 🖼️

N/A.

IDE

N/A; command-line reproduction

Operating system

Windows.

Version

.NET SDK 10.0.401.

Device

N/A.

DynamicData Version

Main 10.0-preview at 3d21360.

Additional information ℹ️

Why this was deferred rather than folded into PR #1191.

The integer path had a direct substitution available. Its moments are held in long, so cross multiplying moves the cancellation sensitive subtraction into integer arithmetic. That option does not transfer to double or float, whose accumulators are already floating point.

Fixing the floating point path is a separate numerical design question. It needs its own design and its own validation, and any credible fix changes returned values for existing callers, which is a compatibility decision worth taking on its own merits rather than attaching to a divisor correction.

Candidate directions, listed for discussion only. None has been evaluated or validated.

  • A centered accumulator tracking a running mean and centered sum of squares. The removal direction needs care here, because the aggregator removes and replaces arbitrary items in arbitrary order, and in the common Welford formulation subtraction is not the exact inverse of addition. More careful formulations exist; one would have to be selected and validated rather than assumed.
  • Accumulating the moments in a wider type for the float selector. This is the cheapest partial mitigation, since float fails earliest.
  • Compensated, Kahan style summation of the moments.
  • Clamping a negative numerator to zero. This converts NaN outcomes into merely inaccurate ones, so it is a mitigation rather than a fix and would mask the underlying loss.

Suggested regression coverage.

The cases above over the float and double selectors, asserting translation invariance: equally spaced values offset far from zero must report their spacing regardless of the offset. The reference value should be computed on centered inputs so the oracle is not itself subject to the cancellation under test, matching the approach already used by SampleStandardDeviation in StdDevFixture.cs.

Related. Copilot review comment 4073251707 on PR #1191 raises the same defect. PR #1191 deliberately stays scoped to divisor placement plus the integer path and defers this here. Issue #1181 covers the divisor placement bug.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions