Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 61 additions & 22 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,13 @@ which parses/introspects/diffs like any other column type; see
handled on the CRUD side (write-side serialization only, no auto-parsing on
read — `mssql-python` doesn't distinguish `json` columns from `nvarchar`).

### Area tagging and filtering
### Area and schema filtering

Two independent, composable ways to narrow which files a command touches:
**area** is an explicit opt-in tag; **schema** is derived automatically from
each file's own SQL.

#### Area tagging

Any migration file or `database/` code file can declare one or more areas by
starting with a `-- area:` comment:
Expand All @@ -76,43 +82,69 @@ first real statement); a `-- area:` comment later in the file doesn't count.
A file with no directive is untagged, and untagged files are treated as
shared/common.

`pgdb compare`, `pgdb migrate check`, and `pgdb migrate apply` all accept:
#### Schema filtering

No tag needed — schema membership is parsed straight out of the SQL itself:
every schema-qualified (or default-schema, when unqualified) table/view/
function/index reference across every statement in the file, DDL or DML
alike, plus any `CREATE SCHEMA name`. A file whose schema(s) can't be
determined (unparseable content, or no table/schema reference in it at all)
is treated the same as an untagged file — always kept.

#### Options

`pgdb compare`, `pgdb migrate check`, `pgdb migrate apply`, `pgdb testdb up`,
and `pgdb testdb reset` all accept:

- `--area NAME` (repeatable) — restrict to files declaring one of the given
areas, **plus every untagged file** (untagged files always stay in scope).
- `--exclude-area NAME` (repeatable) — drop files declaring one of the given
areas; untagged files are never dropped by this.
- `--schema NAME` (repeatable) — restrict to files referencing one of the
given schemas, **plus every file with no detectable schema reference**.
- `--exclude-schema NAME` (repeatable) — drop files referencing one of the
given schemas; files with no detectable reference are never dropped.

Both can be combined; a file matching both an included and an excluded area
is excluded. Passing neither option applies no filtering (the default,
unchanged behavior).
All four can be combined — a file must pass every filter it's subject to (an
area match doesn't excuse a schema mismatch, and vice versa), and a file
matching both an included and an excluded value on the same axis is
excluded. Passing none of them applies no filtering (the default, unchanged
behavior).

```bash
pgdb migrate apply path/to/database/_migration_scripts --url ... --area billing
pgdb compare path/to/database/ --url ... --exclude-area reporting
pgdb migrate check path/to/database/_migration_scripts --url ... --schema billing --exclude-schema reporting
pgdb testdb up --schema billing
```

`compare`'s default report (no `--report-extra-db`) only checks that the
filtered scripts exist correctly in the DB, so it composes safely with area
filtering. Passing `--report-extra-db` together with an area filter also
reports every DB object outside the filtered area(s) as "missing in
scripts" — since the live database has no concept of areas, only the
scripts side is filtered — so treat that combination's "missing in scripts"
results with that in mind (the CLI prints a warning when you combine them).

`pgdb fetch-missing` deliberately has **no** `--area`/`--exclude-area`: it
diffs the full database against scripts to find genuinely untracked
objects, so narrowing the scripts side by area would make every object
tracked only under a different area look "missing" too — and `--write`
would then reconstruct a duplicate file for something that already exists.

`pgdevkit.areas` exposes the same logic for scripting:
and schema filtering. Passing `--report-extra-db` together with either kind
of filter also reports every DB object outside the filtered area(s)/
schema(s) as "missing in scripts" — since the live database has no concept
of areas, and isn't itself filtered by `--schema` either — only the scripts
side is filtered — so treat that combination's "missing in scripts" results
with that in mind (the CLI prints a warning when you combine them).

`pgdb fetch-missing` deliberately has **no** `--area`/`--exclude-area` (or
`--schema`/`--exclude-schema`): it diffs the full database against scripts
to find genuinely untracked objects, so narrowing the scripts side would
make every object tracked under a different area/schema look "missing" too
— and `--write` would then reconstruct a duplicate file for something that
already exists.

`pgdevkit.areas` exposes the tag-filtering logic for scripting:
`parse_areas`/`file_areas` read a file's declared areas, and
`area_allowed`/`filter_by_area` apply the `only`/`exclude` semantics above.
`pgdevkit.migrate.list_migration_files`/`pending_migrations` and
`pgdevkit.parser.parse_directory` take the same `areas`/`exclude_areas`
keyword arguments (`pgdevkit.fetch_missing.find_missing_objects` doesn't,
for the reason above).
`pgdevkit.schemas` exposes the equivalent for schema filtering:
`sql_schemas`/`file_schemas` detect a file's referenced schemas, and
`schema_allowed`/`filter_by_schema` apply the same `only`/`exclude`
semantics. `pgdevkit.migrate.list_migration_files`/`pending_migrations` and
`pgdevkit.parser.parse_directory` take both pairs of keyword arguments
(`areas`/`exclude_areas` and `schemas`/`exclude_schemas`);
`pgdevkit.fetch_missing.find_missing_objects` takes neither, for the reason
above.

## `pgdb testdb`

Expand Down Expand Up @@ -144,6 +176,13 @@ def ensure_test_postgres():

CLI: `pgdb testdb up|reset|run-sql|status|shell|clean`.

`up`/`reset` accept `--area`/`--exclude-area` and `--schema`/`--exclude-schema`
(see "Area and schema filtering" above) to scope which `database/` files get
applied — e.g. `pgdb testdb up --schema billing` for a test DB with only the
`billing` schema's tables/views/functions, without waiting on the rest of the
project's schema to apply. `ensure_testdb`/`reset_testdb` take the same
keyword arguments when called from Python (e.g. from a pytest fixture).

Container connection defaults (`localhost:54322`, `postgres`/`testpwd`) can
be overridden with `PGDEVKIT_TESTDB_HOST`, `PGDEVKIT_TESTDB_PORT`,
`PGDEVKIT_TESTDB_USER`, `PGDEVKIT_TESTDB_PASSWORD`. Before touching the
Expand Down
82 changes: 68 additions & 14 deletions pgdevkit/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,9 +29,21 @@
_EXCLUDE_AREA_OPTION = typer.Option(
[], "--exclude-area", help="Skip files declaring this area (repeatable); untagged files are never excluded"
)
_SCHEMA_OPTION = typer.Option(
[],
"--schema",
help="Restrict to files referencing this DB schema (repeatable); "
"files with no detectable schema reference always stay in scope",
)
_EXCLUDE_SCHEMA_OPTION = typer.Option(
[],
"--exclude-schema",
help="Skip files referencing this DB schema (repeatable); "
"files with no detectable schema reference are never excluded",
)


def _as_area_set(values: list[str]) -> frozenset[str] | None:
def _as_set(values: list[str]) -> frozenset[str] | None:
return frozenset(values) if values else None

testdb_app = typer.Typer(name="testdb", help="Manage the shared local Postgres test container")
Expand Down Expand Up @@ -59,17 +71,20 @@ def compare(
dialect: str = typer.Option("postgres", "--dialect", help="postgres (default) or mssql"),
area: list[str] = _AREA_OPTION,
exclude_area: list[str] = _EXCLUDE_AREA_OPTION,
schema: list[str] = _SCHEMA_OPTION,
exclude_schema: list[str] = _EXCLUDE_SCHEMA_OPTION,
scripts_dir: Path = typer.Argument(..., help="Directory containing SQL scripts"),
) -> None:
"""Compare SQL scripts to a live database and report differences."""
if not scripts_dir.is_dir():
err_console.print(f"[red]Error:[/red] {scripts_dir} is not a directory")
raise typer.Exit(2)
if report_extra_db and (area or exclude_area):
if report_extra_db and (area or exclude_area or schema or exclude_schema):
console.print(
"[yellow]⚠[/yellow] --report-extra-db with --area/--exclude-area will report every DB object "
"outside the filtered area(s) as \"missing in scripts\", since the live database has no concept "
"of areas — only the scripts side is filtered."
"[yellow]⚠[/yellow] --report-extra-db with --area/--exclude-area/--schema/--exclude-schema will "
"report every DB object outside the filtered area(s)/schema(s) as \"missing in scripts\", since the "
"live database has no concept of areas — and isn't itself filtered by --schema either — only the "
"scripts side is filtered."
)

try:
Expand All @@ -91,7 +106,12 @@ def compare(

with console.status("Parsing SQL scripts..."):
scripts_schema = parse_directory(
scripts_dir, dialect=backend.dialect, areas=_as_area_set(area), exclude_areas=_as_area_set(exclude_area)
scripts_dir,
dialect=backend.dialect,
areas=_as_set(area),
exclude_areas=_as_set(exclude_area),
schemas=_as_set(schema),
exclude_schemas=_as_set(exclude_schema),
)

with console.status("Introspecting database..."):
Expand Down Expand Up @@ -188,17 +208,33 @@ def fetch_missing(


@testdb_app.command("up")
def testdb_up() -> None:
def testdb_up(
area: list[str] = _AREA_OPTION,
exclude_area: list[str] = _EXCLUDE_AREA_OPTION,
schema: list[str] = _SCHEMA_OPTION,
exclude_schema: list[str] = _EXCLUDE_SCHEMA_OPTION,
) -> None:
"""Ensure the container is running, the workspace DB exists, and schema is applied."""
testdb.ensure_testdb()
testdb.ensure_testdb(
areas=_as_set(area), exclude_areas=_as_set(exclude_area), schemas=_as_set(schema),
exclude_schemas=_as_set(exclude_schema),
)
info = testdb.status()
console.print(f"[green]Test DB ready:[/green] {info['database']} ({info['dsn']})")


@testdb_app.command("reset")
def testdb_reset() -> None:
def testdb_reset(
area: list[str] = _AREA_OPTION,
exclude_area: list[str] = _EXCLUDE_AREA_OPTION,
schema: list[str] = _SCHEMA_OPTION,
exclude_schema: list[str] = _EXCLUDE_SCHEMA_OPTION,
) -> None:
"""Drop and recreate only this workspace's database, then reapply schema + seed data."""
testdb.reset_testdb()
testdb.reset_testdb(
areas=_as_set(area), exclude_areas=_as_set(exclude_area), schemas=_as_set(schema),
exclude_schemas=_as_set(exclude_schema),
)
info = testdb.status()
console.print(f"[green]Test DB reset:[/green] {info['database']}")

Expand Down Expand Up @@ -272,6 +308,8 @@ def migrate_check(
),
area: list[str] = _AREA_OPTION,
exclude_area: list[str] = _EXCLUDE_AREA_OPTION,
schema: list[str] = _SCHEMA_OPTION,
exclude_schema: list[str] = _EXCLUDE_SCHEMA_OPTION,
) -> None:
"""List which migration files under migrations_dir are applied vs. pending."""
if not migrations_dir.is_dir():
Expand All @@ -281,7 +319,11 @@ def migrate_check(
conninfo = build_conninfo(url, entra_user)
tracking_table = tracking_table or migrate.default_tracking_table(migrations_dir)
local_files = migrate.list_migration_files(
migrations_dir, areas=_as_area_set(area), exclude_areas=_as_area_set(exclude_area)
migrations_dir,
areas=_as_set(area),
exclude_areas=_as_set(exclude_area),
schemas=_as_set(schema),
exclude_schemas=_as_set(exclude_schema),
)
try:
applied = migrate.applied_migrations(conninfo, tracking_table)
Expand Down Expand Up @@ -323,6 +365,8 @@ def migrate_apply(
yes: bool = typer.Option(False, "--yes", "-y", help="Skip the confirm-target prompt"),
area: list[str] = _AREA_OPTION,
exclude_area: list[str] = _EXCLUDE_AREA_OPTION,
schema: list[str] = _SCHEMA_OPTION,
exclude_schema: list[str] = _EXCLUDE_SCHEMA_OPTION,
) -> None:
"""Apply pending migration files, in filename order, tracking each in tracking_table."""
if not migrations_dir.is_dir():
Expand All @@ -331,7 +375,8 @@ def migrate_apply(

conninfo = build_conninfo(url, entra_user)
tracking_table = tracking_table or migrate.default_tracking_table(migrations_dir)
areas, exclude_areas = _as_area_set(area), _as_area_set(exclude_area)
areas, exclude_areas = _as_set(area), _as_set(exclude_area)
schemas, exclude_schemas = _as_set(schema), _as_set(exclude_schema)
target_desc = url.rsplit("@", 1)[-1] if "@" in url else url
if not yes:
typer.confirm(f"About to run migrations against {target_desc}. Continue?", abort=True)
Expand All @@ -341,13 +386,22 @@ def migrate_apply(
else:
try:
targets = migrate.pending_migrations(
migrations_dir, conninfo, tracking_table, areas=areas, exclude_areas=exclude_areas
migrations_dir,
conninfo,
tracking_table,
areas=areas,
exclude_areas=exclude_areas,
schemas=schemas,
exclude_schemas=exclude_schemas,
)
except migrate.TrackingTableMissing:
err_console.print(
f"[yellow]⚠[/yellow] {tracking_table} not found — treating every migration as pending"
)
targets = migrate.list_migration_files(migrations_dir, areas=areas, exclude_areas=exclude_areas)
targets = migrate.list_migration_files(
migrations_dir, areas=areas, exclude_areas=exclude_areas, schemas=schemas,
exclude_schemas=exclude_schemas,
)

if not targets:
console.print("No pending migrations.")
Expand Down
10 changes: 10 additions & 0 deletions pgdevkit/dialect.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,16 @@
}


# Schemas that hold system catalog views/tables, never a file any project
# using pgdevkit manages -- a reference to one (e.g. an idempotency guard
# querying it, or a `SELECT ... FROM information_schema/pg_catalog/sys ...`)
# is never a real schema-membership or cross-file-dependency signal. Shared
# by `schemas.py` (schema-reference filtering) and `testdb/schema.py`
# (dependency-safe apply ordering), which both walk the same sqlglot Table
# nodes for a related-but-different purpose.
SYSTEM_SCHEMAS = {"pg_catalog", "information_schema", "sys"}


@dataclass(frozen=True)
class Dialect:
"""A thin wrapper around a sqlglot dialect name plus the handful of
Expand Down
Loading
Loading