I build data platforms and intelligent products that connect reliable pipelines, clear architecture, useful interfaces, and operational reality.
|
From issue to verified pull request. K0 coordinates durable repository-role workflows. It authorizes work, snapshots the committed workflow, invokes AI or human roles, validates evidence, follows explicit graph edges, and keeps human review at the closing boundary.
|
Exact accounting with an auditable agent boundary. An independently installable accounting kernel for immutable double-entry records, evidence-backed proposals, human approval, receipt-safe posting, reproducible statements, and typed access for applications and agents.
|
Data engineering is the foundation of my work: ingestion and transformation, orchestration, data contracts, catalog identities, lineage, quality, governed access, and the cloud platforms that keep those capabilities dependable.
I am interested in agents that can do meaningful work without making authority ambiguous. That means explicit roles, typed tools, least-privilege access, durable state, resumable execution, observable events, verifiable evidence, and human checkpoints where judgment or irreversible action begins.
My recent work connects agent orchestration with real platform boundaries: metadata discovery, document consultation, controlled read-only SQL, delegated specialists, asynchronous tool execution, usage accounting, and failure-aware client events.
- OpenLineage - merged contributions for Spark adaptive plans, catalog identities, lineage logging, and Spark UI / Databricks configuration.
- Marquez - merged SQL pagination and JobDao work.
- K0 - durable human-and-agent software-delivery workflows; coming to open source.
- Atomic Accounting - evidence-backed, agent-accessible accounting infrastructure; coming to open source.
- From PoC to viable product: industrializing a Big Data project - DevFest Lille, 2022
- OpenLineage in Action: A Catalyst for Data Observability - OpenLineage × Apache Kafka at Confluent London, 2024
- Former Master 2 Big Data & AI lecturer at ESEM Paris
- Delivered professional Airflow, Spark, and dbt training across +15 sessions to +150 participants
| Area | Technologies and practices |
|---|---|
| Languages and formats | Python · SQL · Scala · TypeScript · JavaScript · Java · C · C++ · Fortran · PHP · Bash · MATLAB · YAML · XML · JSON · PDF processing |
| Data engineering | PySpark · Apache Spark · Hadoop · Hive · Airflow · MWAA · Kafka · dbt · Talend for Big Data · Dataiku · Google Cloud Data Fusion · Azure Data Factory · ETL/ELT · DataFrames · materialized views |
| Data platforms and governance | Databricks · Unity Catalog · AWS Glue · OpenLineage · Marquez · metadata platforms · data contracts · data quality · catalog identity reconciliation · dataset discovery · lineage · Gold/Silver/Insight models |
| Cloud and storage | AWS · Azure · GCP · S3 · EMR · Redshift · EKS · Lambda · SQS · SageMaker · DynamoDB · Azure Data Lake · BigQuery · PostgreSQL · MongoDB · Algolia |
| Applications and APIs | React · Next.js · FastAPI · Flask · Django · OpenAPI · DocAPI · REST APIs · SSO · Shopify · storefronts · back-office systems · web dashboards · Android |
| Generative AI and agents | LLMs · AWS Bedrock · LangGraph · LangChain · LiteLLM · MCP · tool calling · multi-agent delegation · document search · controlled SQL access · token accounting · asynchronous execution · delegated event streaming |
| Machine learning and analytics | NumPy · pandas · SciPy · scikit-learn · TensorFlow · TensorFlow C++ · NLTK · NetworkX · BigQuery ML · Azure ML · clustering · collaborative filtering · recommendation systems · neural networks · Power BI · Google Data Studio · seaborn · Matplotlib |
| Scientific computing | MPI · graph algorithms · routing and spectrum assignment · linear and nonlinear optimization · parallel computing |
| Platform engineering | Docker · Kubernetes · EKS · FluxCD · Kustomize · nginx · IAM · GitHub Actions · Jenkins · ECR · CodeArtifact · Nexus · Code Factory · Vault · Checkmarx · CloudWatch · CI/CD · observability · logging · alerting |
| Testing and reliability | TDD · unit testing · integration testing · pytest · AsyncMock · asyncio · Vitest · Testing Library · typed interfaces · immutable records · receipt-safe retries · reproducible builds |
| Delivery and collaboration | Git · GitHub · Bitbucket · Jira · Confluence · Linux · technical specifications · architecture reviews · evidence-backed delivery |
| Business systems and sources | SAP · Lisa · Shopify · e-commerce and order workflows · pricing and revenue-management data · customer analytics |





