Skip to content
DM

Projects

Things I have built

Internal platforms described by what they do and what they run on, and open source work you can read yourself.

01

At Stryker

Five systems the data engineering and data science organization runs on, each built end to end.

Data platform application with Almanac AI

Stryker

Internal data platform deployed as a Databricks App, bringing platform monitoring, an AI data assistant, governance workflows, in-browser data editors, and notifications into one place for engineers, admins, analysts, and business users.

  • Built Almanac, an AI data assistant on Claude through Databricks Model Serving. A streaming tool-use loop explores Unity Catalog, writes and runs SQL, and answers with the source tables cited, and a planner splits complex questions into parallel sub-questions
  • Grounded answers with retrieval over embedded table and column comments, pipeline code, SQL, and Power BI model files, plus a knowledge graph that derives join paths from Power BI relationships and verifies each join against live data before trusting it
  • Made the assistant safe on production data with a sqlglot validator that allows only single read-only SELECTs, AI-written PySpark run under a read-only identity on serverless jobs, and a result cache that avoids re-billing the warehouse
  • Exposed Almanac to other applications as an OAuth-secured API with a consumer registry, quotas, rate limits, and resumable server-sent event streaming, and bridged it into Microsoft Teams to answer mentions and direct messages as the person asking
  • Built registry-driven monitoring for app health, data readiness, table freshness with statistical stall detection, Power BI refreshes, DBU usage, SQL warehouse health, API traffic, GitHub vulnerabilities, and CI/CD, with daily rollups that turn a 160 million row audit view into sub-second pages
  • Delivered every alert to a stored history, live in-app toasts, and Microsoft Teams Adaptive Cards with per-user subscriptions, and built access-request approvals that apply GRANT, REVOKE, and Entra group changes on approval with drift enforcement and a full audit log
  • Grew the codebase to roughly 500 API endpoints, 42 pages, and 180 thousand lines across Python and TypeScript, nearly all of it authored solo
  • Python
  • FastAPI
  • Granian
  • Next.js
  • React
  • TypeScript
  • Tailwind CSS
  • Databricks
  • Unity Catalog
  • Delta Lake
  • Claude
  • Microsoft Graph
  • Entra ID

Central CI/CD and repository governance

Stryker

Organization-wide CI/CD and repository governance built from nothing as one central repository of reusable GitHub Actions workflows, consumed by more than 28 data engineering and data science repositories.

  • Built one pipeline that detects Python, Node.js, Rust, and Docker in each repository and runs only the matching lint, build, test, and scan jobs, replacing several older workflows and MegaLinter
  • Designed an AI pull request reviewer on Claude through Databricks Model Serving that completes a structured checklist, posts prioritized inline findings with suggested fixes, re-checks every finding with a second model call before posting, answers developer replies on review threads, and blocks merges on unresolved high-priority issues through a required status check
  • Put supply-chain security into every pull request with Semgrep, Secretlint, Bandit, pip-audit, Dependency Review with license blocking, CycloneDX SBOMs, Trivy container scans, and build provenance attestations, reported to GitHub Code Scanning as SARIF
  • Automated repository configuration as code with a daily Python tool that enforces branch rulesets, required status checks, team access profiles, and encrypted secrets across more than 30 repositories, handling rate limits and updating rulesets in place so protection is never lost mid-run
  • Automated pull request governance with size labels, per-path approval gates, reviewer assignment from the history of the changed files, a guard that closes edits to centrally managed files, and Dependabot auto-merge for minor and patch updates once every required check passes
  • GitHub Actions
  • GitHub REST and GraphQL
  • Python
  • JavaScript
  • Bash
  • Claude
  • Ruff
  • Semgrep
  • Trivy
  • Docker

Common functions library

Stryker

Shared Python library for Databricks pipelines, Azure authentication, and Microsoft 365 and Power BI integrations, distributed as a versioned wheel to downstream data teams.

  • Built a config-driven medallion pipeline framework where three dictionaries for read, transform, and write settings compile into complete incremental PySpark pipelines across Bronze, Silver, and Gold layers, removing hand-written Spark and SQL
  • Wrote the transform language covering joins, unions, aggregations, pivots, window functions, deduplication, CASE logic, safe division, and inline cast and text markers, with reusable builders for rolling windows, year-over-year growth, aging buckets, and fiscal calendars
  • Implemented incremental loading on Delta Change Data Feed with watermark and full-refresh fallbacks, merge-on-key writes with upstream deletion scans guarded against empty sources, and automatic schema evolution inside the same overwrite
  • Standardized cleansing, hash-based primary keys, data quality thresholds, a per-run metrics table, Delta conflict retries with exponential backoff, and serverless-aware performance with skew handling, broadcast joins, and scratch-table spooling
  • Built the authentication and connection layer with Entra ID client-credentials tokens, Key Vault secrets, Microsoft Graph clients for SharePoint, mail, and users, a Power BI client, and a rate-limited Databricks REST client covering jobs, clusters, warehouses, Unity Catalog, and MLflow
  • Maintained roughly 45 thousand lines across 130 modules with unit tests, Ruff, Bandit, and Semgrep in CI, and a wheel built and versioned on every pull request
  • Python
  • PySpark
  • Delta Lake
  • Databricks
  • Unity Catalog
  • Azure Key Vault
  • Entra ID
  • Microsoft Graph
  • Power BI
  • GitHub Actions

Customer master harmonization

Stryker

Production entity resolution on Databricks Serverless, built on custom matching algorithms, that unifies customer records from five ERP and CRM systems into one governed customer master, replacing a discontinued third-party vendor while keeping legacy IDs for traceability.

  • Designed the matching engine from scratch with seven parallel blocking strategies and key-frequency capping that cut a naive comparison space of tens of billions of pairs down to tractable candidate sets, then scored candidates with custom matching algorithms, a Fellegi-Sunter log-odds model blended with twenty hand-built similarity features across names, phonetics, addresses, and geography
  • Built name and address normalization that strips company suffixes while keeping healthcare descriptors, standardizes street types and directionals, extracts former and alternate names, and detects garbage addresses so scoring weight shifts from address to name
  • Assigned stable customer master IDs that survive reruns through sequential allocation, address and phonetic grouping, anchor matching against history, and a precedence chain where manual review always beats automation
  • Added an LLM matcher for records fuzzy matching could not resolve, combining Databricks-hosted embeddings, Spark-side blocking with a top-k cosine rerank, and a Claude judge on Databricks Model Serving that returns schema-validated JSON with a confidence score and rationale
  • Wrote the judge's rubric around the dominant false-positive mode, co-located healthcare tenants sharing one address, so identity must come from the name, and cut input-token cost with prompt caching while pre-normalized embeddings made each similarity check a single dot product
  • Made AI matching crash-safe and idempotent with batched staging of every judgment, resumable runs that skip conclusive results and retry transient failures, adaptive rate pacing, dry runs, and a single re-runnable Delta merge that re-derives downstream tables without new LLM spend
  • Solved Spark Connect and serverless constraints with a central temp-table manager that materializes lineage, cleans up in a finally block, and sweeps stale tables, across roughly 21 thousand lines of modular Python with pytest suites, Databricks Asset Bundles, and GitHub Actions CI
  • Python
  • PySpark
  • Delta Lake
  • Unity Catalog
  • Databricks Serverless
  • Databricks Model Serving
  • Claude
  • Custom matching algorithms
  • Databricks Asset Bundles
  • GitHub Actions

Engineering wiki

Stryker

Team knowledge wiki rebuilt from MkDocs into a custom Next.js application deployed as a Databricks App, covering onboarding, CI/CD, data engineering, and data science standards across more than 110 pages.

  • Built the markdown pipeline on unified, remark, and rehype with admonitions, GitHub alert syntax, KaTeX math, Shiki code highlighting, and heading anchors, pre-rendering every page at build time with redirects from the old URLs
  • Added full-text search on an in-memory Orama index generated at build time, served through an API route and a command palette
  • Designed documentation federation that pulls docs from five source repositories on a signal, daily, or on demand, and opens pull requests into the protected main branch so each team keeps ownership of its own pages
  • Recorded per-page last-updated dates, created dates, and contributors in a single git log pass and persisted them so a build without git history still shows them
  • Built in-page viewers for draw.io, Mermaid, PDF, Word, and Excel files, and shipped with a strict Content Security Policy, HSTS, and the standard hardening headers
  • Next.js
  • React
  • TypeScript
  • unified
  • Orama
  • Databricks Apps
  • GitHub Actions