Skip to content

Adopt strict ruff, mypy and typos, and reformat the codebase - #128

Merged
zeevmoney merged 40 commits into
mainfrom
per-16222/strict-tooling
Oct 2, 2026
Merged

zeevmoney merged 40 commits into
mainfrom
per-16222/strict-tooling

Conversation

@zeevmoney

@zeevmoney zeevmoney commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Linear issues

  • Fixes PER-16222: adopt strict ruff, mypy and typos, and reformat and fix the codebase to pass them.
  • Part of PER-16336: the SDK upgrade guide, section 10 (packaging and tooling).
  • Related to PER-15743: the backend now refuses an invite whose role belongs to another resource than its target, and the invites e2e test follows it.

Also supersedes Dependabot's #134 (ruff 0.16.8) and #135 (mypy 2.3.1). Based on main after #127 (uv migration).

Summary

  • Strict linting, formatting, type checking and spell checking for the whole repository, run from the versions locked in uv.lock.
  • The codebase is reformatted and fixed to pass them.
  • No change to the public API, apart from one regenerated model (ApproveMessage) and two bug fixes, listed under Behaviour.

Why

  • The SDK was linted by ruff 0.6.9 with pyupgrade disabled, and type-checked by a non-strict mypy 1.11.2.
  • mypy did not check the tests.
  • Its pre-commit environment could not see aiohttp or loguru, and under pydantic 2 it checked the SDK's pydantic.v1 models against the v2 API.
  • The pre-commit hooks installed their own ruff and mypy, so the versions in uv.lock were not the ones that ran.

What changed

Tooling

  • ruff 0.16.8:
    • select = ["ALL"], line length 100, Google docstring convention.
    • Every ignore has a one-line reason in pyproject.toml.
    • The generated permit/api/models.py is excluded from lint and format.
  • mypy 2.3.1, strict:
    • Adds warn_unreachable and 11 more error codes.
    • Covers permit/, tests/, scripts/, skills/ and .github/scripts/.
    • Runs under both pydantic majors, with the pydantic.v1.mypy plugin.
  • typos 1.50.2: spell checking.
  • Versions: the newest releases outside the repo's 7-day cooldown.
  • pre-commit:
    • ruff, ruff-format, mypy and typos run as repo: local hooks through uv run --locked, so uv.lock is the single source of their versions.
    • The mypy hook also runs when pyproject.toml or uv.lock changes.
    • pre-commit-hooks v6.0.0 is pinned by SHA, and check-shebang-scripts-are-executable is added.
  • pytest: strict = true and filterwarnings = ["error"] for the SDK suite, skills/tests and .github/scripts.
  • Dependabot:
    • A new pre-commit ecosystem entry.
    • ruff, mypy and typos get their own uv group, so a lint-rule change can't hold up runtime dependency bumps.
  • CONTRIBUTING.md documents the checks and how the hooks sync .venv.

CI

  • pre-commit.yml runs every hook, then mypy again on the pydantic-1 lane.
  • The compatibility job drops its -W flag, because the pytest config now turns warnings into errors.
  • e2e PDP start:
    • The job waits up to 300 s of elapsed time for the PDP to report healthy, up from 180 s. Each health probe times out after 5 s.
    • In CI, the PDP has taken 63–154 s to start.
  • e2e failure log: when the job fails, the PDP log it prints drops the once-a-second health checks. It keeps the first Health check failed: horizon line, which says why the PDP was not healthy.

Code

  • Mechanical changes, each in its own commit: ruff format, then ruff's safe fixes, such as absolute imports and pyupgrade.
  • Hand-written changes:
    • Google docstrings on the public API.
    • Type annotations, with if TYPE_CHECKING: imports of pydantic.v1 at the version-conditional import sites. The runtime branches are unchanged.
    • @overload where a return type depends on an argument.
    • Explicit re-exports in permit/__init__.py.
  • Test fixes:
    • test_envs checked an environment created in a different project, and a reused variable made its finally block raise AttributeError and hide the real failure. mypy found both.
    • The invites e2e test now gives its invites a role of the invited resource, which the backend requires (PER-15743).

Behaviour

  • Public API unchanged:
    • dir(permit), from permit import *, every @validate_arguments model and every public signature are identical to main on both pydantic majors.
    • ModelListInput[X] keeps the runtime annotation List[X] that 3.0.0 has, and a regression test pins it.
    • Runtime-visible aliases (Context, AuthorizedUsersDict, IncEx, User, Resource) and the bare-dict pydantic fields are kept, each with a justified suppression. permit.api.base.TData is kept because 3.0.0 exports it.
  • One regenerated model: ApproveMessage's only field is now detail instead of message, to match the live API spec. No SDK method returns this model. The schema drift check passes.
  • Two fixes, each with tests that fail on main:
    • jsonable_encoder(Decimal("NaN")), and the same for sNaN and Infinity, raises a TypeError that names the value, because JSON has no such values. It used to raise an unrelated TypeError from comparing the Decimal's exponent with 0. The exception type is unchanged.
    • import permit no longer emits a DeprecationWarning. PermitConnectionError subclassing the deprecated PermitException used to emit one. Code that instantiates or subclasses PermitException still gets the warning.

How it was tested

  • Local checks:
    • pre-commit passes on all files.
    • mypy strict reports no issues in 89 files, under pydantic 1.10.26 and 2.13.5.
    • uv lock --check passes.
    • actionlint and zizmor are clean.
  • Offline suite: 305 passed, 3 skipped, 0 warnings, on both pydantic lanes.
  • Other suites:
    • skills/tests: 86 passed, 1 skipped.
    • .github/scripts: 108 passed.
    • tests/test_typing_surface.py: 3 passed.
  • e2e in CI, against a scratch environment and a local PDP, with warnings as errors: 322 passed, 7 skipped, on both pydantic lanes.
  • Build: the wheel file list, sdist file list and METADATA are identical to main.
  • Runtime surface: a snapshot of every permit module, pydantic field and @validate_arguments model matches main on both majors. The only differences are ApproveMessage and the two fixes.
  • Fail-without-fix:
    • reverting either fix fails its new tests;
    • returning list[X] from ModelListInput fails its regression test.
  • Review: an independent reviewer compared the whole runtime surface with main and approved. Their findings are fixed in this branch.

Manual test plan

  1. With uv 0.12.17 or later, run uv sync, then uv run pre-commit run --all-files. All hooks pass.
  2. Run uv run --group pydantic-v1 mypy. It reports no issues.
  3. Run uv run pytest -m "not e2e". It reports 305 passed and 3 skipped, with warnings as errors.
  4. Run uv run python -W error::DeprecationWarning -c "import permit". It exits 0.
  5. Run uv run python -c "from decimal import Decimal; from permit.api.encoders import jsonable_encoder; jsonable_encoder(Decimal('NaN'))". It raises TypeError: Decimal('NaN') is not JSON serializable: JSON has no NaN or Infinity.

Blast radius

  • Code: the formatter touches every Python file. The runtime surface matches main, apart from ApproveMessage and the two fixes.
  • Contributors: the hooks run the locked tools, so uv sync is all that's needed. A commit re-syncs .venv to the default groups.

Scope and size

  • permit/: +2,610 / -2,007, mostly formatting, docstrings and annotations.
  • Tests: +1,372 / -662.
  • skills/, scripts/ and .github/: +1,480 / -589.
  • Config and lock: pyproject.toml +185 / -45, .pre-commit-config.yaml +42 / -18, uv.lock +299 / -39.
  • Where to look: the format and safe-fix commits are mechanical. Review the config commit, the hand-written code commits, the two fixes and the CI changes.

🤖 Generated with Claude Code

zeevmoney and others added 18 commits September 21, 2026 17:19
The resolved dependency tree was clean, but the published `>=` floors let a
consumer install versions carrying 34 known advisories. Because this package
ships open ranges with no lockfile, the floor is the real exposure -- so the
scan covers both the current resolution and the lowest versions the specs
permit.

Dependency fixes:
- aiohttp >=3.14.3 (clears 32 advisories, incl. CVE-2026-69244, an
  out-of-bounds heap read in the HTTP response parser this client exercises
  on every call)
- pydantic >=1.10.13 (CVE-2024-3772, EmailStr ReDoS; the SDK uses EmailStr)
- werkzeug >=3.1.6, pytest >=9.0.3
- drop httpx: never imported, and the only path by which h11
  (CVE-2025-43859, CRITICAL) and anyio entered the tree
- drop zipp and aioresponses: both unused, and aioresponses 0.7.9 is
  incompatible with aiohttp 3.14.3
- python_requires >=3.10; the declared >=3.8 was already unachievable

Gates:
- Trivy over three trees (runtime ceiling, runtime floor, dev), sticky PR
  comment, blocking on fixable HIGH/CRITICAL only
- release split into build -> scan -> publish, so publish is unreachable
  unless the scan passed
- weekly cron posting the findings themselves to Slack, not just a verdict
- Dependabot with cooldowns and versioning-strategy: increase
- delete release.yml, which raced python-sdk-publish.yml on every release
- existing workflows hardened: 48 zizmor findings (12 high) to zero

Also fixes 10 minor SDK bugs with 33 offline regression tests. Nine major
correctness bugs found along the way are tracked in PER-16174 rather than
changed here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…endent

pytest_httpserver's `httpserver` fixture is session-scoped: the first test
that requests it binds the one shared server for the entire run. The address
override lived in test_rbac_e2e.py, so it only applied when that module
happened to touch the fixture first.

Adding tests/test_offline_regressions.py broke that assumption -- it sorts
earlier, claimed the session server on a random port, and test_api_timeout
and test_pdp_timeout then failed against their hardcoded localhost:9999 with
"Cannot connect to host".

Moving the fixture to conftest.py makes the address apply session-wide and
removes the latent ordering dependency, which any future test using
httpserver would otherwise have tripped over too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
get, get_by_key, update and delete all interpolate their argument straight
into the path, and the backend validates it with
validate_resource_instance_ident(instance_id, allow_uuids=True) -- a bare
instance key is rejected with a 422, not accepted. The docstrings said "the
key of the resource instance", which sends callers straight into that error.

Wording matches what bulk_delete already documented correctly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bumps to 3.0.0 and fixes the nine major bugs tracked in PER-16174, so the
eight permanently-xfail tests can assert for real.

Sync client (permit/utils/sync.py, permit/sync.py):
- SyncClass is now idempotent. It was inherited, so a subclass re-wrapped
  methods its base had already converted, giving async_to_sync(async_to_sync(f));
  all 21 deprecated-facade methods raised "a coroutine was expected" before
  issuing a request.
- Coroutine detection uses inspect.iscoroutinefunction and unwraps
  functools/validate_arguments wrappers, instead of assuming every object whose
  class is named "function" is async.
- permit.sync.Permit now overrides authorized_users, get_user_permissions and
  filter_objects, which were inherited as `async def` over a synchronous
  enforcer and returned un-awaitable coroutines.

Enforcement (permit/enforcement/):
- parse_obj_as is imported through the pydantic v1/v2 guard the rest of the
  package uses; authorized_users() could not return at all under pydantic v2.
- bulk_check honours a per-check context and filter_objects forwards the
  caller's context. It was silently dropped, so context-dependent ABAC
  evaluated against {} and could return the wrong subset.
- UserInput accepts snake_case as well as the camelCase aliases; first_name
  and last_name were silently discarded from every check.

Serialization (permit/api/base.py):
- dict and list bodies go through the encoder, so nested datetime/UUID/Enum
  no longer dies inside aiohttp.
- exclude_none is dropped, so an explicitly-set None is transmitted as null
  and an update can clear a field. exclude_unset still omits untouched fields.

Facts proxy (permit/api/tenants.py):
- tenants bulk operations addressed the PDP's users endpoint.

tests/endpoints/test_bulk_operations.py asserted that a tenant role assignment
outlives the user who owns it; deleting the user removes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The un-xfailed tests all run against one shared environment and were fighting
each other: fixed keys (admin, viewer on the built-in __tenant resource), a
shared resource urn, assertions on global object counts, and teardown that
called pytest.fail on a 404 so "already deleted by another test" turned a
passing test red. Several also leaked every object they created.

Each test now derives its keys from tests/utils.unique_key, asserts against
its own objects rather than environment-wide counts, tears down in a finally
via handle_cleanup_error, and polls with a bounded retry where it waits for a
fact to reach the PDP. Verified by running twice in a row against a
deliberately dirty local environment.

test.yml starts the PDP as a step rather than a service container. A service
container is created before the first step runs, so it could only be given the
long-lived PROJECT_API_KEY while the tests authenticate with the per-run
scratch environment key. The PDP rejected every decision with a 403, which is
why the ReBAC and RBAC decision tests could never pass.

That 403 also surfaced as "cannot connect to the PDP container": the enforcer
read error bodies with response.json(), and the PDP sends auth rejections as
plain text, so ContentTypeError -- an aiohttp.ClientError -- was caught by the
connectivity handler and the real status was lost. Error bodies are now read
without assuming JSON, and the message names the status and body.

tests/test_abac_pdp.py's three cloud-PDP tests now skip with a reason instead
of failing: as CI is configured they never reach the cloud PDP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The PDP reports 503 on /healthy until its horizon component finishes pulling
config and a policy bundle. Waiting for it immediately after docker run made
that bootstrap serial with the job; one leg was ready in 29s and the other
still was not at 60s. The wait now happens after dependency installation, so
the bootstrap overlaps with it, with a 180s ceiling.

Changing an ABAC condition set makes the policy generator recompile the
environment's rego and redistribute the bundle, which is much slower than the
fact sync RBAC uses. test_abac_e2e timed out at 90s against the real cloud PDP;
raised to 300s. The poll returns as soon as the rule lands, so a healthy run is
no slower.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
setup.py used a bare find_packages(), which ships a TOP-LEVEL `tests` package
into every consumer's site-packages where it shadows their own `tests` module.
Verified against the published permit==2.8.3, which does exactly that. Now
excluded, along with `harness`.

permit.pdp_api never passed a timeout to its HTTP client, so the documented
pdp_timeout was silently ignored on every permit.pdp_api.* call while the
enforcer honoured it. It also duplicated ClientConfig and pagination_params
verbatim from permit.api.base; it imports them now.

Removed, none of which had a single caller in permit/, tests/ or harness/:
  set_if_not_none (enforcer), OpaResult and the JWT alias (interfaces),
  ApiKeyLevel (a self-declared deprecated alias of ApiKeyAccessLevel),
  LoginAsErrorMessages (never compared against or returned), and three unused
  TypeVars in the PDP base module.

_model_dump was defined identically in both arms of the pydantic version
split; hoisted to one definition. Its `mode` parameter stays and stays
ignored on purpose -- it absorbs a v2-style argument that pydantic v1's
.dict() would reject.

Repo cruft: .isort.cfg (isort is not run; ruff's I rules are), uv.lock (a
three-line stub declaring requires-python >=3.14, contradicting setup.py),
the Makefile publish target (a second release path that bypasses the gated
build -> scan -> publish workflow) and a .DEFAULT_GOAL pointing at a help
target that did not exist. .gitignore's .DS_Store rule was inert because of
an inline comment.

Dependencies: dropped pytest-mock (no test uses it) and pytest-cov (coverage
is never requested, including in CI). Corrected the werkzeug comment -- it is
now a direct test import, not just a pytest_httpserver transitive.

Also dropped two references to .trivyignore, which audit-deps.sh deliberately
disables with --ignorefile /dev/null, so both were advertising a suppression
mechanism that does not work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
The condition sets and rule this test creates never reach the PDP's policy
bundle, so the decision it waits for never becomes true. The PDP says so in
the debug.abac payload the SDK already logs: ~90s of no_matching_usersets
with "known usersets: ['rules']" (the empty-package placeholder), then one
bundle carrying only the condition sets autogenerated by the resource and
role creates ten seconds earlier, then nothing for the remaining 300s. The
data channel stayed healthy throughout.

The pipeline is event-driven with no polling fallback (the default scope is
created with poll_updates=False and batching drains rather than waits), so
this is a stall, not slowness, and no timeout makes it pass. Skipped rather
than xfailed so it reports honestly instead of looking like coverage.

Only the three decision assertions are skipped. Everything above them still
runs against the real control plane -- condition set and rule create, type
round-trip, paginated list, filtered list, permission-format assertion -- and
so does the teardown, because pytest.Skipped derives from BaseException and
escapes the test's except Exception.

Ruled out as causes: resource_id passed as .hex (the generator keys on the
resource key, never the id), inline check attributes (they win the
object.union_n in the generated rego and the PDP echoed them back), and a
missing setup step.

No other test is exposed: condition_set_changes.py is the only policy
synchronizer handler that generates rego, so RBAC and ReBAC decisions resolve
against data.* on the fact channel, and this is the only test that touches
condition sets.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
resource_relations.list() declared List[RelationRead], but the route is
declared response_model=PaginatedResult[RelationRead], so against current
backend main the call raised "ValidationError: value is not a valid list" --
the method was unusable. It now returns PaginatedResultRelationRead; callers
read .data. BREAKING, and in the 3.0.0 notes.

(That change was written earlier and swept into the previous commit by a
bare `git add -A`; this records what it actually is.)

Two docstrings corrected against the backend, both of which sent callers into
a confusing error:

- resource_roles.assign_permissions/remove_permissions said permissions are
  <resourceKey:actionKey>. A resource role is scoped to its own resource, so
  each entry is a BARE action key. Passing the qualified form makes the server
  read the whole string as an action key and reject it with a 404 naming
  '<resource>:<resource>:<action>' -- a doubled prefix that reads like the SDK
  concatenated wrongly, when it is the server quoting what it was given.

- role_assignments.list(resource_instance_key=...) takes a
  `resource_type:instance_key` ident or an instance uuid, never a bare key.

Regression tests pin the exact wire strings on both pydantic majors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
Every remaining CI failure was one cause: HTTP 429 on a cleanup call. Enabling
the eight previously-xfail tests and giving each its own objects made the suite
create and tear down far more than before, and teardown is where the burst
lands -- one leg reported 3 failed and 2 teardown errors, the other 7 failed,
all of them 429 on a delete.

handle_cleanup_error now tolerates 429 alongside 404, for the same reason 404
is tolerated: neither leaves the test's assertions in doubt. A throttled delete
leaks an object, and CI deletes the whole scratch environment afterwards, so it
is reclaimed. Any other status still fails the test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
The previous commit tolerated 429 during teardown. That was wrong in a way the
next CI run made obvious: a tolerated DELETE leaves the object alive, so the
assert-it-is-gone check that follows failed with "DID NOT RAISE
PermitApiError". The tolerance manufactured a worse failure than the one it
hid. 429 is no longer tolerated.

It was also the wrong layer. The run after showed 429 arriving in test BODIES
as well -- test_rebac_e2e, test_sync_client and test_user_invites_complete_e2e
all failed mid-test -- so cleanup was never the whole problem. The suite runs
against one environment on a shared cloud project and now creates and tears
down considerably more than it used to, which exceeds the burst limit. The
eight tests that were xfail until this branch had been swallowing these 429s
all along.

conftest wraps the SDK's five HTTP verbs for the test session only, retrying a
429 with exponential backoff so the call actually succeeds. The SDK is
untouched: adding implicit retries to a published client would be a behaviour
change callers did not ask for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
Six attempts (~63s of backoff) still ran out on one teardown, leaving CI at
1 failed / 102 passed. Raised to nine, which caps a single call at roughly two
minutes of waiting and exits the moment it succeeds.

Also honours the server's Retry-After when it sends one, and adds jitter to
the exponential fallback so concurrent callers do not retry in lockstep and
re-trip the limit together.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
bulk_check() reads each query's context with .get(), so a query without
one is valid at run time, but the TypedDict declared the key as required
and mypy rejected every bulk_check([{"user", "action", "resource"}]) call.
TypedDict comes from typing_extensions so NotRequired is honoured on 3.10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
pyproject.toml now carries the PEP 621 metadata setup.py declared, built
with uv_build; dev tools move to a PEP 735 group and both pydantic lanes
become conflicting groups, so every CI lane installs from the committed
uv.lock. setup.py, requirements*.txt, MANIFEST.in, pytest.ini and the
Makefile are gone; contributor docs move to CONTRIBUTING.md.

CI installs with uv sync --locked; the publish job stamps the version
with uv version, builds with uv build --no-sources on a checksum-verified
uv, and keeps its build -> scan -> publish gating and PyPI token auth.
The audit compiles its three trees from pyproject.toml with --no-sources
and fails if the dev group did not resolve. uv is pinned once, by
[tool.uv] required-version, with a 7-day exclude-newer cooldown;
Dependabot uses the uv ecosystem and a uv-lock hook stops drift.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
ruff 0.16.7 with select = ["ALL"] minus justified ignores, line length
100 and Google docstrings; mypy 2.3.1 strict over permit/, tests/ and
.github/scripts on both pydantic majors, with TYPE_CHECKING branches so
the v1 models type-check as v1 under pydantic 2. ruff, mypy and typos
run as local pre-commit hooks from uv.lock (uv run --locked), external
hooks are SHA-pinned, pytest runs strict with warnings as errors, and
Dependabot covers pre-commit with lint tools grouped apart from runtime
floors.

No public API or behaviour change; runtime-visible aliases, bare-dict
fields and the star-import surface are kept identical. Three bugs the
stricter checks exposed are fixed with regression tests: decimal_encoder
crashed on NaN/Infinity, a pre-release pydantic version crashed
import permit, and import permit raised under -W error because
PermitConnectionError subclasses the deprecated PermitException.

py.typed is deliberately not shipped yet (PER-16231).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F6b4ERDYYZ8NRTv1zJYxx2
@linear-code

linear-code Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

PER-16222

PER-16336

PER-15743

Base automatically changed from per-16221/uv-migration to main September 29, 2026 22:49
zeevmoney and others added 9 commits September 30, 2026 01:55
Main gained the 3.0.0 SDK fixes (#126) and the final uv migration
(#127) after this branch was cut. The branch reformatted and strictly
typed the pre-3.0.0 code, so the merge conflicted in most files.

The tree is reset to main's tree here, so the tooling, formatting and
typing changes can be re-applied on top of the 3.0.0 code in separate
commits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ruff 0.16.8 with `select = ["ALL"]`, line length 100, Google docstring
  convention and docstring-code-format. Every ignore is justified in
  pyproject.toml. The generated permit/api/models.py and the migration
  skill's sample apps stay out of lint and format (force-exclude), and the
  migration scanner is held to Python 3.8 syntax.
- mypy 2.3.1 `strict`, plus warn_unreachable and extra error codes, over
  every Python file but the generated models and the sample apps, with the
  pydantic.v1 mypy plugin on both pydantic majors.
- typos 1.50.2 checks spelling.
- pytest runs with `strict = true`.
- ruff, ruff-format, mypy and typos are `repo: local` pre-commit hooks
  running `uv run --locked`, so uv.lock is the only source of their
  versions. pre-commit-hooks v6.0.0 is pinned by SHA and adds
  check-shebang-scripts-are-executable.
- CI type-checks once more under pydantic 1.
- Dependabot gets a pre-commit ecosystem entry, and ruff, mypy and typos
  a group of their own in the uv entry.

The code is reformatted and fixed in the following commits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The deprecation-warning test expected each warning on the line after its
helper's `def`, which stops being true once the formatter wraps the
helper's signature. Read the line of the helper's one statement from its
syntax tree instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mechanical: `ruff format` with the configuration from the previous
commit. The sync stub generator lays out permit/_sync_types.pyi the way
ruff format does at a given line length, so its LINE_LENGTH moves to 100
and the stub is regenerated; the result is what ruff format produces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The generator only resolved names brought in with `from module import`.
An annotation such as `builtins.list[str]`, which a class that defines a
`list` method needs, refers to a module imported whole with
`import builtins`; the generator now emits that import in the stub, in
the order ruff's isort rules use. The committed stub does not change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mechanical: `ruff check --fix` (safe fixes only), then `ruff format`, and
the sync stub regenerated from the fixed classes. Most of it is PEP 585
and 604 annotations, docstring layout, else-after-return and sorted
imports.

Three rewrites would have changed runtime objects, so those sites keep
their spelling with a noqa that says why:
- `Context` and `AuthorizedUsersDict` are public aliases, so they stay
  `typing.Dict` generics rather than becoming builtin ones.
- `UserInput.attributes`, `ResourceInput.attributes` and
  `ResourceInput.context` stay `typing.Dict`: pydantic v1 validates a
  `typing.Dict` value into a copy but keeps the caller's object for a
  bare `dict`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The safe fixes rewrote the `IncEx` alias in permit/api/encoders.py with
builtin generics (`set[int]`, `dict[str, Any]`), which changes the
runtime object the alias names. Restore the `typing` generics it had,
with a noqa, as for `Context` and `AuthorizedUsersDict`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SDK now passes `ruff check` and strict mypy under both pydantic
majors. Most of it follows the approach of the original PR-128 commit:
- absolute imports, return and parameter annotations, `ParamSpec` on
  `handle_client_error`, `@overload` on `delete()`, and Google
  docstrings on the public API (the `Sync*` runtime classes included);
- `TModel` is no longer bound to `BaseModel`, since list endpoints
  parse into `list[Model]`, and the unused `TData` is removed;
- equivalent rewrites the rules ask for: HTTPStatus constants, messages
  assigned before `raise`, `input` renamed where it shadowed the builtin.

Runtime-visible spellings are kept, with a suppression that says why:
the `User`, `Resource` and `_UserSyncInput` aliases, the bare-`dict`
pydantic fields, `PermitConnectionError`'s deprecated base, and the
positional signatures of four `list()` methods (PLR0917).

`UserInput.attributes`, `ResourceInput.attributes` and
`ResourceInput.context` become `dict[Any, Any] | None`, which pydantic
v1 validates exactly like the `Optional[Dict]` they were: into a copy
of the caller's dict. A new test fails if they ever become a bare
`dict`, which keeps the caller's object and lets the tenant the SDK
adds leak into it.

Tooling that goes with it:
- PLC0414 is off: `import X as X` is the explicit re-export strict
  mypy needs, and the SDK uses it for the blocking classes in
  permit/_sync_types.pyi.
- The stub's copied docstrings are allowed (PYI021).
- The stub generator accepts a docstring in the `Sync*` runtime classes,
  and the stub is regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The org-level environment test created each environment under `project`,
the variable its project loop left behind, which is unbound when the loop
does not run and otherwise names whichever project came last. The
assertions that follow check `projects[0]`. Create the environments in
`projects[0]` too. mypy reports the old line as possibly undefined.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@zeevmoney
zeevmoney marked this pull request as ready for review September 30, 2026 00:30
Copilot AI balanced review requested due to automatic review settings September 30, 2026 00:30

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

In CI the PDP's /healthy has taken 63-154s to return 200. It stays 503
until the scratch environment's first policy bundle and data arrive,
and until then the PDP restarts its policy service about once a minute.
On main after #127 the pydantic-2 job ran past the 180s limit on three
attempts, while the same job passed in every PR run.

Wait up to 300s. The step still prints how long the PDP took. On
failure, drop the PDP's once-a-second health-check lines before taking
the tail of its log, so the policy and data fetches that explain a slow
start are no longer cut off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 30, 2026 15:13

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@EliMoshkovich EliMoshkovich left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deep review done. I compared runtime behaviour between main (040a744) and this branch under pydantic 1.10.18/1.10.26 and 2.4.2/2.13.5 on py3.10 and py3.11.

Verified clean:

  • dir(permit), from permit import * and every pydantic model's fields are identical on both majors.
  • Every @validate_arguments signature and validation model is the same, with 11 sample inputs coerced identically through every field.
  • The enforcer's request bodies, headers, paths, error types and messages are byte-identical against a local PDP stub, for statuses 200/403/500/501 and connection refused.
  • jsonable_encoder output is identical.
  • SyncClass wraps the same methods, and _sync_types.pyi regenerates with zero diff.
  • Exception hierarchy: MRO, except PermitException, pickling and warnings are all unchanged, apart from the intended import-time silence.
  • Tooling runs clean locally: pre-commit on all files, mypy strict on both pydantic lanes, 305 passed + 3 skipped offline on both majors, skills/tests 86 + 1 skipped, and .github/scripts 108 passed.
  • scan.py output is byte-identical to main on py3.8 and py3.9.
  • uv.lock changes only dev tools. No runtime floor or ceiling moved.
  • Test changes were AST-compared with main; none are weakened, and test_error_response::test_api_error is stricter.

Inline findings: one should-fix, one risk to decide on, and two CI nits. None of them is a blocker. Details are inline.

Comment thread permit/utils/model_input.py Outdated
Comment thread tests/test_offline_regressions.py Outdated
Comment thread pyproject.toml
Comment thread .github/workflows/test.yml Outdated
Comment thread .github/workflows/test.yml Outdated

@EliMoshkovich EliMoshkovich left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving: nothing I found blocks the merge. Please do the ModelListInput should-fix (a 2-line change: return List[model] and restore the test assertion) before merging, so the "identical runtime surface" claim holds. The filterwarnings point and the two test.yml nits are your call, and fine as follow-ups.

zeevmoney and others added 3 commits October 2, 2026 20:28
Ruff's UP006 fix changed ModelListInput[X] at runtime from
typing.List[X] to list[X]. The two do not compare equal, so
get_type_hints() on the bulk methods' undecorated functions returned a
different annotation than in 3.0.0. Validation was unaffected.

Return typing.List[X] again, and restore the regression test's
assertion that the annotation equals List[UserCreate].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The wait loop's curl had no timeout, so a PDP that accepted the
connection and never answered could hold the step indefinitely. Each
probe now gives up after 5s.

On failure, the PDP log filter dropped every "Health check failed:
horizon" line, including the one that says why the PDP never became
healthy. Keep the first such line; the later repeats and the GET
/health requests are still dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With each probe allowed 5s, 300 tries could take far longer than the
300s the error message reports. The loop now stops once 300s have
passed, whatever the probes took.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 17:32

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

zeevmoney and others added 2 commits October 2, 2026 21:07
The invites target a resource instance, and the API now refuses to
approve an invite whose role belongs to another resource (PER-15743).
The test gave them a tenant role; it now creates a role on the invited
resource and deletes it before the resource.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The live spec renamed ApproveMessage's only field from message to
detail, so the schema drift check failed on it. No SDK method returns
this model. The class is the generator's output, copied unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 18:09

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@zeevmoney zeevmoney changed the title Adopt strict ruff and mypy, reformat and fix the codebase Adopt strict ruff, mypy and typos, and reformat the codebase Oct 2, 2026
@zeevmoney
zeevmoney merged commit ef80ae2 into main Oct 2, 2026
29 of 30 checks passed
@zeevmoney
zeevmoney deleted the per-16222/strict-tooling branch October 2, 2026 18:24
zeevmoney added a commit that referenced this pull request Oct 2, 2026
main is the squash of #128, whose commits this branch already has, so
the tree is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants