Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 23 additions & 1 deletion .github/workflows/python-test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -31,4 +31,26 @@ jobs:
- name: ty check
run: uv run ty check pgdevkit
- name: Test with pytest
run: uv run -m pytest --capture=tee-sys --maxfail=3 tests
run: uv run -m pytest --capture=tee-sys --maxfail=3 -m "not mssql" tests

mssql-test:
# Separate job (not folded into `build`) so a live SQL Server -- a much
# larger image and slower cold start than the Postgres container --
# never slows down or blocks the fast, always-run Postgres suite above.
# mssql-python bundles its own ODBC driver, so unlike pyodbc this needs
# no system driver package install here.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
submodules: "recursive"
- name: Set up Python 3.14
uses: actions/setup-python@v5
with:
python-version: "3.14"
- name: Install uv
run: curl -LsSf https://astral.sh/uv/install.sh | sh
- name: Install project dependencies
run: uv sync --all-extras --all-groups
- name: Run MSSQL-only tests
run: uv run -m pytest --capture=tee-sys --maxfail=3 -m mssql tests
45 changes: 41 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,22 @@ pgdb compare --url postgresql://instance-abc.database.azuredatabricks.net:5432/d
(`--url`'s own user/password, if any, are discarded and replaced — `--entra-user`
plus the fetched token become the connection's actual credentials.)

### MSSQL

`pgdb compare`/`pgdb fetch-missing` default to Postgres. Pass `--dialect mssql`
to compare against a SQL Server database instead:

```bash
pgdb compare --dialect mssql --url "Server=host,1433;Database=db;UID=user;PWD=pass" path/to/database/
```

Requires the `mssql` extra: `pip install pgdevkit[mssql]` (pulls in
[mssql-python](https://github.com/microsoft/mssql-python), which bundles its
own driver — no system ODBC driver install needed). MSSQL has no composite
type, native enum, or first-class JSONB column type, so those areas of a
`database/` tree don't have a direct equivalent on this backend — see
`docs/database-layout.md`.

## `pgdb testdb`

Manages a single shared, Podman-backed Postgres container for local tests
Expand Down Expand Up @@ -89,12 +105,29 @@ The role named by `PGDEVKIT_TESTDB_USER` must exist and match your OS user
(`CREATE ROLE <user> SUPERUSER LOGIN;`) and `pg_hba.conf` must allow `peer`
auth for local connections (Debian/Ubuntu Postgres ships this by default).

### MSSQL

Add `engine = "mssql"` to `[tool.pgdevkit]` (or set
`PGDEVKIT_TESTDB_ENGINE=mssql` for a one-off run) to manage a shared SQL
Server container instead of Postgres — same one-container-per-machine,
one-database-per-workspace model. Requires the `mssql` extra (see above).

Container defaults (`localhost:14330`, `sa`/a generated complexity-valid
password) can be overridden with `PGDEVKIT_TESTDB_MSSQL_HOST`, `_PORT`,
`_USER`, `_PASSWORD`, `_IMAGE`, `_MEMORY_LIMIT_MB`. The container only
bootstraps the `sa` login — additional logins are a known limitation.
`pgdb testdb shell` execs into
[`sqlcmd`](https://github.com/microsoft/go-sqlcmd) (an external prerequisite,
the same category as `psql` for the Postgres path) rather than a Python
REPL.

## `pgdevkit.db` — helpers for application code

Install with the `db` extra: `pip install pgdevkit[db]`.

- **`PostgresTableModel`** — a `pydantic.BaseModel` base class for models
that map 1:1 to a table row. Implement `get_table_name()` (returns
- **`TableModel`** (formerly `PostgresTableModel`, still importable under
that name) — a `pydantic.BaseModel` base class for models that map 1:1 to
a table row, for either engine. Implement `get_table_name()` (returns
`(schema, table)`) and `get_primary_key()` on each model.
- **`PgPool`** — an async connection pool keyed off
`{env_prefix}HOST/PORT/DB/USER/PASSWORD` env vars. Call `await pool.open()`
Expand All @@ -107,8 +140,12 @@ Install with the `db` extra: `pip install pgdevkit[db]`.
- **CRUD functions** — `pg_retrieve`, `pg_retrieve_many`, `pg_insert`,
`pg_insert_many`, `pg_update`, `pg_update_dict`, `pg_upsert`,
`pg_upsert_dict`, `pg_upsert_many`, `pg_upsert_many_dict`, `pg_delete`,
`pg_delete_dict` — typed (`PostgresTableModel`-based) or dict-based CRUD
against a table, built on `psycopg` for safe identifier/value handling.
`pg_delete_dict` — typed (`TableModel`-based) or dict-based CRUD against a
table, built on `psycopg` for safe identifier/value handling. The `mssql`
extra provides an `mssql_*`-prefixed mirror of the same functions in
`pgdevkit.db.mssql_crud`, built on `mssql-python` (`MERGE`-based upsert,
`OUTPUT` instead of `RETURNING`) — MSSQL has no composite/enum/JSONB
equivalent, so `complex_helper` is always `None` on that path.
- **`SqlLoader`** — loads and caches `.sql` files from
`{root}/<topic>/<name>.sql`, for keeping hand-written queries out of
Python source.
Expand Down
15 changes: 15 additions & 0 deletions docs/database-layout.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,21 @@ comment on column dim.user.is_active is 'False once a user is soft-deleted; keep

---

## MSSQL projects (`engine = "mssql"`)

Everything above is engine-agnostic *as a folder/apply-order convention*,
with two exceptions:

- `types/` (custom types / enums) has no direct T-SQL equivalent -- MSSQL
has neither a native enum type nor composite types, so a `CREATE TYPE ...
AS ENUM`/composite `.sql` file is a Postgres-only construct. `pgdb compare`
reports every such object as missing on an MSSQL database (correctly --
it genuinely doesn't exist there), rather than erroring.
- T-SQL scripts conventionally separate batches with a standalone `GO` line
(an `sqlcmd`/SSMS scripting convention, not valid inside a single
driver `execute()` call). `pgdb testdb` splits on these automatically when
applying a file; hand-written `.sql` files may use `GO` freely.

## Backfilling untracked objects

If a table, scalar function, or table function was created directly on the
Expand Down
21 changes: 21 additions & 0 deletions pgdevkit/backends/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
from __future__ import annotations

from ..dialect import Dialect, resolve_dialect
from .base import Backend
from .mssql import MssqlBackend
from .postgres import PostgresBackend

_REGISTRY: dict[str, Backend] = {
"postgres": PostgresBackend(),
"mssql": MssqlBackend(),
}


def get_backend(dialect: str | Dialect = "postgres") -> Backend:
"""Look up the `Backend` for a dialect name (or an already-resolved
`Dialect`). Defaults to postgres."""
resolved = resolve_dialect(dialect)
return _REGISTRY[resolved.name]


__all__ = ["Backend", "MssqlBackend", "PostgresBackend", "get_backend"]
27 changes: 27 additions & 0 deletions pgdevkit/backends/base.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
from __future__ import annotations

from typing import Any, Callable, Protocol

from ..dialect import Dialect
from ..models import DatabaseSchema


class Backend(Protocol):
"""Introspection + a couple of engine facts, behind one interface.

CRUD is deliberately NOT part of this protocol -- psycopg's
`AsyncConnection` and an MSSQL driver's connection type are unrelated,
so a unified `backend.retrieve()`/`backend.insert()` surface would force
existing Postgres callers to go through a new indirection just to keep
working. Callers that want CRUD import `pgdevkit.db.crud`'s `pg_*`
functions or `pgdevkit.db.mssql_crud`'s `mssql_*` functions directly,
exactly as `db/crud.py`'s functions are imported today."""

dialect: Dialect

def introspect(self, conninfo: str) -> DatabaseSchema: ...

def complex_helper_factory(self) -> Callable[..., Any] | None:
"""A `ComplexHelper`-like factory for composite/enum/JSONB columns,
or None when the engine has no equivalent (MSSQL)."""
...
27 changes: 27 additions & 0 deletions pgdevkit/backends/mssql.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
from __future__ import annotations

from typing import Any, Callable

from ..dialect import MSSQL, Dialect
from ..models import DatabaseSchema


class MssqlBackend:
dialect: Dialect = MSSQL

def introspect(self, conninfo: str) -> DatabaseSchema:
# Imported lazily so `import pgdevkit.backends` (and thus
# `pgdevkit.cli`) doesn't require mssql-python/the mssql extra to
# be installed unless a caller actually asks for the mssql backend.
from ..mssql_introspect import introspect_mssql_db

return introspect_mssql_db(conninfo)

def complex_helper_factory(self) -> Callable[..., Any] | None:
# MSSQL has no composite type, native enum, or first-class JSONB
# column type -- there is nothing for a ComplexHelper to adapt.
# Every `complex_helper` parameter in db/crud.py (and its
# db/mssql_crud.py counterpart) is already Optional, so callers on
# this backend simply pass/receive None and every complex-type
# branch takes its existing no-op path.
return None
19 changes: 19 additions & 0 deletions pgdevkit/backends/postgres.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
from __future__ import annotations

from typing import Any, Callable

from ..dialect import POSTGRES, Dialect
from ..introspect import introspect_db
from ..models import DatabaseSchema


class PostgresBackend:
dialect: Dialect = POSTGRES

def introspect(self, conninfo: str) -> DatabaseSchema:
return introspect_db(conninfo)

def complex_helper_factory(self) -> Callable[..., Any] | None:
from ..db.complex_types import ComplexHelper

return ComplexHelper
23 changes: 16 additions & 7 deletions pgdevkit/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,10 @@
from rich import box

from . import testdb
from .backends import get_backend
from .connection import build_conninfo
from .diff import DiffKind, compute_diff
from .fetch_missing import SUBFOLDER, find_missing_objects, layer_folder_for, reconstruct_ddl
from .introspect import introspect_db
from .parser import parse_directory

app = typer.Typer(name="pgdb", help="PostgreSQL database schema tools")
Expand All @@ -37,9 +37,10 @@ def compare(
None, "--databricks-instance", help="Lakebase instance name (required for Lakebase hosts)"
),
report_extra_db: bool = typer.Option(False, "--report-extra-db", help="Report objects in DB but not in scripts"),
dialect: str = typer.Option("postgres", "--dialect", help="postgres (default) or mssql"),
scripts_dir: Path = typer.Argument(..., help="Directory containing SQL scripts"),
) -> None:
"""Compare SQL scripts to a live PostgreSQL database and report differences."""
"""Compare SQL scripts to a live database and report differences."""
if not scripts_dir.is_dir():
err_console.print(f"[red]Error:[/red] {scripts_dir} is not a directory")
raise typer.Exit(2)
Expand All @@ -55,13 +56,19 @@ def compare(
err_console.print(f"[red]Error:[/red] {e}")
raise typer.Exit(2)

try:
backend = get_backend(dialect)
except ValueError as e:
err_console.print(f"[red]Error:[/red] {e}")
raise typer.Exit(2)

with console.status("Parsing SQL scripts..."):
scripts_schema = parse_directory(scripts_dir)
scripts_schema = parse_directory(scripts_dir, dialect=backend.dialect)

with console.status("Introspecting database..."):
db_schema = introspect_db(conninfo)
db_schema = backend.introspect(conninfo)

diffs = compute_diff(scripts_schema, db_schema, report_extra_db=report_extra_db)
diffs = compute_diff(scripts_schema, db_schema, report_extra_db=report_extra_db, dialect=backend.dialect)

if not diffs:
console.print("[green]No differences found.[/green]")
Expand Down Expand Up @@ -204,8 +211,10 @@ def testdb_status() -> None:

@testdb_app.command("shell")
def testdb_shell() -> None:
"""Drop into psql against this workspace's database."""
os.execvp("psql", ["psql", testdb.dsn_for()])
"""Drop into an interactive shell (psql, or sqlcmd for MSSQL) against
this workspace's database."""
binary, argv = testdb.shell_argv()
os.execvp(binary, argv)


@testdb_app.command("clean")
Expand Down
3 changes: 2 additions & 1 deletion pgdevkit/db/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,13 +17,14 @@
pg_upsert_many_dict,
)
from .loader import SqlLoader
from .model import PostgresTableModel
from .model import PostgresTableModel, TableModel

__all__ = [
"ComplexHelper",
"PgPool",
"PostgresTableModel",
"SqlLoader",
"TableModel",
"pg_delete",
"pg_delete_dict",
"pg_insert",
Expand Down
11 changes: 9 additions & 2 deletions pgdevkit/db/model.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,10 @@
from pydantic import BaseModel


class PostgresTableModel(BaseModel, ABC):
"""Base class for models that map 1:1 to a database table/row.
class TableModel(BaseModel, ABC):
"""Base class for models that map 1:1 to a database table/row (any
engine -- schema/table naming is equally meaningful for Postgres and
MSSQL, this base class was never actually Postgres-specific).

Models representing partial results (joins, aggregations, projections)
should extend `pydantic.BaseModel` directly instead."""
Expand All @@ -21,3 +23,8 @@ def get_table_name() -> tuple[str, str]:
@abstractmethod
def get_primary_key() -> Sequence[str]:
"""Return the primary key column name(s)."""


# Backward-compat alias -- this class was named PostgresTableModel before
# MSSQL support existed; kept so existing imports keep working unchanged.
PostgresTableModel = TableModel
Loading
Loading