Type Cardinality Calculator for your projects.
The primary goal of this project is to improve the quality of Scala code,
whether it is written by a person or generated by a large language model. Both
do their best work when the types they are given are small: a type with few
possible values leaves little room to invent a case that cannot happen, or to
miss one that can. Types with large cardinality (or, worse, unbounded ones like
String or List) push the burden of correctness onto ad-hoc runtime checks
and guesswork.
scala-cardinality measures how many values a type can hold, so you can see —
and shrink — the state space that has to be reasoned about. Smaller, more
precise types give tighter constraints and a better chance of producing code
that compiles and behaves correctly.
For LLMs in particular, this matters more than ever: the types are part of the context a model reasons over, so a smaller state space means fewer invalid states to hallucinate and a much better shot at correct first-try output.
This project makes that feedback loop concrete: parse Scala source and it reports the cardinality of the types it finds.
Started from a blog post around why this makes sense to do: https://medium.com/swlh/type-system-saving-55db2ca062ff
And a later presentation around this @ Scala Love: https://slides.com/rodolfohansen/keep-your-types-small
Key references motivating this work:
- Counting type inhabitants (Alex Knvl, 2018): the type arithmetic condensed in docs/type-arithmetic.md
- The Hidden Powers of Total Program Cardinality
- The Hidden Powers of Total Program Cardinality (2 of 2)
The v1 implementation plan tracks the agreed counting contract, the latest recorded eo baseline, and the remaining work. Report snapshots below are historical.
The calculator is an early work in progress. It parses Scala source with scalameta and counts products, tagged sums, exponentials, powersets and primitives, following the type arithmetic condensed in docs/type-arithmetic.md, and reports the cardinality of the definitions a source introduces — classes, enums, modules, type aliases and opaque types, nested definitions included.
Sizes are polynomials over three tiers — a·ε₀ + b·ω + n — added componentwise, so two
countable alternatives stay 2ω (Either[String, String]) instead of collapsing into one
ω, while products and powers stay deliberately coarse ((String, String) is ω).
Counter.source sums the value spaces of the definitions in a source;
Counter.sourceSignature sums what their members declare instead, so an object with two
String => String methods and one String field is 2ε₀ + ω.
Scala 3 extension methods contribute too: each method's domain includes its receiver,
the extension group's using clauses, and its own parameters. Methods in an extension
group are added separately; the receiver is not also counted as a field.
Givens contribute their declared instance type once, including named and anonymous aliases,
abstract declarations, and template-based instances. Parameterized givens count as factories
over their parameter domain; methods inside a given's implementation are not counted again.
Recursive types are counted as fixed points: each source is solved as a system of equations by
Kleene iteration from the empty type, so a strict case class Loop(next: Loop) comes out empty
while Peano-style Nat is countably infinite (ω), and forward references resolve exactly. A
recursive component that keeps growing is widened to the tier it has grown into, one of the
solver's approximations: only a component that can reach itself is widened, so however long a
chain of forward references is, it settles exactly, and an ε₀ payload is never demoted.
Laziness is what separates the least fixed point from the greatest: a cycle through a position
the constructor never demands — a hole (=> X, () => X), a strict Option[X] field, or a
function field D => X — also counts its infinite values, so
case class Stream(h: Boolean, t: => Stream) reaches ω. At completion of each lazy type,
the finite-value μ estimate and infinite contribution are combined, then normalized to one ε₀
if that tier is present, otherwise one ω if present; purely finite totals stay unchanged.
Thus LazyList[String] and Stream[String] are ε₀, conaturals and LazyList[Boolean]
are ω, and LazyList[Nothing] and a pure lazy self-wrapper remain 1.
Finite and infinite families still exist, but this approximation deliberately discards their
breakdown at the lazy-type boundary; it is not exact cardinal or ordinal arithmetic.
Enclosing sums and source/signature aggregates are not normalized:
Either[LazyList[String], LazyList[String]] is 2ε₀.
Any branching cycle saturates at ω (the §4 finite-program reading: one program per
unfolding). Open work: coinduction blocked by a strict
self argument or a deeply nested mention (kept at a sound under-count), parametric counts and
sealed hierarchies spanning files — pinned as pending targets in the test suite.
docs/ holds the pages that get published as the documentation site — currently
type arithmetic, the counting rules the calculator targets and
the tests that pin them down. The README stays the entry point for the repository itself.
sbt siteRender # renders docs/ into target/site
python3 -m http.server -d target/site 8000The renderer is Laika with its Helium theme — the same
engine the sister project eo uses.
Two pieces of eo's pipeline are missing here, both because this build runs on sbt 2:
sbt-typelevel-site (which wraps Laika for sbt) and sbt-mdoc (which compiles
scala mdoc fences) have no sbt 2 builds, and mdoc would additionally collide with this
project's scalameta_3 dependency (scalameta_2.13 and scalameta_3 share package
names). So the render step is a task in build.sbt plus
project/SiteRenderer.scala, and the numbers shown in the
pages are pinned by the test suite instead
of being compiled from the pages.
CI renders the site on every pull request (ci.yml, "Documentation site" job, artifact
docs-site), and deploy-site.yml publishes it to
Cloudflare Pages — a preview per pull request, production on v* tags — once the
CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID secrets exist together with the
scala-cardinality-docs Pages project. Until then that workflow skips with a notice.
The build has two modules:
core— the calculator itself (Counter, theSizealgebra,Report, the method analysis): a pure library with no sbt types, tested with specs2.plugin—sbt-cardinality, an sbt 2 plugin that reports on the build it is added to, and on any Scala sources it is pointed at. It is tested end-to-end with sbt's scripted framework (sbt plugin/scripted).
Everything builds with the Scala version sbt 2.0.x itself runs on (3.8.4),
because the plugin — and core, which it loads — must be binary-loadable
inside sbt, and Scala 3 binary compatibility is backward only.
sbt cardinalityReport measures every definition the project's Compile
sources introduce — nested definitions included — and logs a report: what each
generic method or constructor can be, then the stored-value estimate over the
definitions, ordered by how many values each holds. The same report is written to
target/cardinality/report.txt (cardinalityReportFile), so a build can keep
it, diff it, or post it as an artifact.
scala-cardinality — 1 source, 10 definitions
stored-value estimates: constructor inputs only; finite capacities are upper bounds
?: unresolved, not a proof of infinity; opaque representations are not singletons
2 unresolved · 4 with more than one value · 2 with one value · 2 abstract
Stored-value estimates (constructor inputs; `?` = unresolved)
? example.Holder[A] class Light.scala:11 unresolved: A
ω example.Succ class Light.scala:17
ω example.Timeline class Light.scala:19
2 example.Custom class Light.scala:5
1 example.Zero object Light.scala:16
— example.Light abstract Light.scala:3
A row carries the number and the reason: ω is a countably infinite value space (a recursive
type solved as a fixed point, a lazy hole that unfolds for ever), ε₀ the tier above it,
2^32 an upper bound on a finite count the calculator tracks in bits, ? a size some name
stopped it from bounding, and — a definition with no cardinality of its own (an abstract
type). A ? row still carries the size the calculator reached, and the summary tells definitions
whose size depends on their own type parameters apart.
The report reads the algebra's sums as they are, coefficients included: Either[String, String]
is ω + ω, and a module with two String => String methods and one String field is
2ε₀ + ω. What the calculator cannot bound, it names: a ? row says both that the number
reached ω or beyond and which type names stopped it from being tighter.
sbt "cardinalityReportOf <path>..." runs the same report over sources the
build does not compile itself — a directory, a single file, or the
-sources.jar a published library ships. That is how a dependency's types get
measured from outside its build:
# eo-core, the `core` module of the sister project `eo`, as published
cs fetch --sources dev.constructive:cats-eo_3:0.16.0
sbt 'cardinalityReportOf <cache>/cats-eo_3-0.16.0-sources.jar'Its run over eo-core 0.16.0 (53 sources, 134 definitions) reads: 7 unresolved, 25
instantiation-dependent, 5 unbounded by an open abstraction, 2 type constructors with no value
space, 3 with more than one value (the two 2^32 array builders and a countable PSVec.Slice),
81 holding a single value (the modules and the sealed parents whose one case is a module),
5 with no values (the X aliases), 6 abstract. A library of generic optics has no small state
spaces to find; what the report says about it is why each row has no number — the type
parameters a generic class leaves open, or the capability an unsealed trait leaves to its
implementations.
The full run is checked in at
docs/baselines/eo-core-0.16.0.txt, with the analyzer
revision, the sources jar's SHA-256, the tool versions and the analysis limits in its header, so
a later run can be diffed against it.
The other number a report gives is the count of canonical pure, total, parametric
implementations a signature admits with everything in scope — what the method can access, capture
or call. def choose[A](x: A, y: A): A has 2 (x or y); the constructor of a case class
Pair[A](x: A, y: A) has 4 ways to build its product. That is a different question from the
stored-value estimate on the data type: Pair[Int] still holds 2^64 values.
ω is a productive cycle: def use[A](x: A, step: A => A): A can return x, step(x),
step(step(x)), … — countably many. A cycle with no starting inhabitant is 0, not ω. An
enclosing value, a callable producer and a product projection all count as captures, and a type
parameter's identity is per binder, so a shadowed A is not an outer A.
What the analysis cannot read is ?, with the reason, and the section header sums the reasons
into the triage list:
Generic method / constructor implementation cardinalities
53 signatures: 41 finite · 3 countably infinite · 9 unresolved
8 unresolved on: unresolved type
1 unresolved on: given environment not resolved
4 example.Pair.<init> Pair[A](left: A, right: A) Light.scala:11 [constructor]
1 example.Accessor.get get[X, A](fa: (X, A)): A Accessor.scala:18 [method]
captures: tupleAccessor
For a direct inherited method, override identity also compares Scala 3 using clauses and
legacy implicit clauses. Contextual arguments remain supplied inputs: a method with
a: A and using fallback: A, returning A, has two choices when no other producer is in
scope. Its matching inherited declaration is not supplied again as a recursive capability.
Parameter names and alpha-renamed method binders do not distinguish slots; nominal parameter
types do. Ordinary/contextual convention mismatches with otherwise matching parameter keys,
dependent results, unknown imports, and additional unsupported
parameter modifiers remain obligations.
This does not implement implicit search or normalize context-bound syntax into explicit clauses.
Transparent first-order aliases are expanded in direct inherited signatures. For example,
type Id[A] = A makes get[A](a: Id[A]): Id[A] the same slot as get[B](b: B): B.
Expansion uses the alias's declaration scope, retains receiver substitutions across shadowing
method binders, lets a nearer declaration or import shadow a farther binder, keeps a qualified
selection distinct from a same-spelled method formal, and never equates distinct nominal types
merely because their shapes agree. Closed, qualified, chained, and tuple aliases are supported.
Recursive or opaque aliases, abstract type members, bounded or higher-kinded alias parameters, and
type-lambda aliases remain guarded; expansion respects the configured type-depth limit, including
nested argument positions.
The refreshed eo-core 0.16.0 run reads 480 signatures: 30 finite, 2 countably infinite, 448 unresolved. The review ledger records the reproduction command, all fifteen numeric changes since the previous baseline, and five independently derived model counts checked against the original archive. The other numeric rows remain provisional. Major overlapping obligations include abstract/member-bearing representations (225 rows), unresolved types (183), polymorphic capability binders (104), inherited override identity (78), and higher-kinded parameters. These are named next steps, not claims about the code.
The build mirrors the sister project
eo: the same formatter,
linter, coverage, mutation-testing and duplicate-detection tools, split between
the always-on gates in .github/workflows/ci.yml
and the heavier reports in
.github/workflows/quality.yml.
| Tool | Purpose | Local command | CI |
|---|---|---|---|
| scalafmt | Formatting | sbt scalafmtAll |
ci.yml, check-only, gating |
| scalafix | Semantic rewrites (unused/organized imports, syntax bans) | sbt scalafixAll |
ci.yml, check-only, gating |
| scoverage | Statement/branch coverage | sbt coverageAll |
ci.yml, gating on a coverage floor |
| scripted | Plugin end-to-end tests | sbt plugin/scripted |
ci.yml, gating |
| stryker4s | Mutation testing | sbt mutationAll |
quality.yml, on PRs, report only |
| CPD (PMD) | Duplicate-code detection | PMD's pmd cpd (see the cpd job) |
ci.yml, gating, duplicates of 25+ tokens |
scripts/check-file-metrics.sh |
Per-file length and comment ratio | bash scripts/check-file-metrics.sh |
ci.yml, gating; warns above 800 lines or 25% comment lines, fails above 1000 or 35% |
| CodeScene | Code Health and hotspots | cs delta |
quality.yml, on PRs, gating once CS_ACCESS_TOKEN is set |
Configuration lives in .scalafmt.conf, .scalafix.conf, stryker4s.conf,
.codescene/custom-quality-gates.json and the thresholds at the top of
scripts/check-file-metrics.sh. Coverage is gated just below the current
baseline (a regression floor, expected to ratchet up); mutation testing never
fails the build and exists to guide test investment.
The Scala 3 compiler options in build.sbt follow the recommendations
from sbt-typelevel-settings,
with the broader -Wunused:all checks and -Werror: warnings
fail both production and test compilation, locally and in CI. On Scala 3.9+
the build enables -opt and -opt-inline:<sources>, limiting bytecode inlining
to the current compilation's sources rather than dependencies. These optimizer
flags are disabled while the plugin and core target sbt's Scala 3.8.4 runtime.
On pull requests each workflow posts its results, passing and failing alike, as
one comment that is edited in place on every push: ci.yml the gates with the
coverage rates, any CPD duplicates and the file-metric outcome, quality.yml
CodeScene delta. Other runs write the same table to the run summary.
sbt scalafmtAll # apply formatting
sbt scalafixAll # apply semantic fixes
sbt coverageAll # tests + coverage report under target/
sbt plugin/scripted # plugin end-to-end tests (fresh sbt per test project)
sbt mutationAll # core mutation report under target/stryker4s-report/Note
sbt 2 keeps a machine-wide task cache ($XDG_CACHE_HOME/sbt, usually
~/.cache/sbt) that clean does not clear. Re-running a gate can therefore
report "no tests to run" and skip the coverage check — a cache hit, not a
failure. Force a cold run with a cache namespace you have not used yet:
sbt -Dsbt.cacheversion=1 coverageAll (reusing a value resolves to the same
entries). Restored instrumentation that outlives its scoverage-data/
directory can also surface as a scoverage FileNotFoundException during a real
test run; a fresh namespace clears that too. (sbt --sbt-cache <dir> needs the
stock sbt launcher rather than the sbt-launch jar, so it is rejected here.) CI is
unaffected: the runner starts with an empty cache.
CodeScene's primary integration is its GitHub App, which reviews pull requests
against the quality gates in .codescene/custom-quality-gates.json. The
quality.yml job adds a CLI cs delta gate on top of that, on pull requests and
pushes to main. It fails on any Code Health finding, runs only when the
CS_ACCESS_TOKEN repository secret is set, and otherwise posts a notice and
skips.