Experience
I build data platforms, LLM services and agentic systems, from implementation to production operations.
Certifications & continued learning
AI & LLMs4
NVIDIANVIDIA · LLMs, transformers, parallelism and diffusionView credentials
Rapid Application Development with Large Language Models (LLMs)
NVIDIA · 2024
Verify credential ↗Model Parallelism: Building and Deploying Large Neural Networks
NVIDIA · 2023
Verify credential ↗Generative AI with Diffusion Models
NVIDIA · 2024
Verify credential ↗Building Transformer-Based Natural Language Processing Applications
NVIDIA · 2024
Data & ML Engineering4
IBM
DeepLearning.AI
Stanford OnlineIBM · DeepLearning.AI · Stanford OnlineView credentials
Data Topology Design
IBM · 2022
Verify credential ↗Machine Learning Engineering for Production (MLOps)
DeepLearning.AI · 2021
Verify credential ↗Neural Networks and Deep Learning
DeepLearning.AI · 2020
Machine Learning
Stanford Online · 2020
Verify credential ↗
Responsible Technology2
Green Software
PennRegulatory compliance and green softwareView credentials
Green Software Practitioner
Green Software Foundation · 2026
Verify credential ↗Regulatory Compliance
University of Pennsylvania · 2020
Verify credential ↗
Professional English1
EF SET · C2EF SET · C2 Proficient · 75/100 · four skillsView credentials
EF SET English Certificate · C2
EF Education First
Verify credential ↗
Translate business requirements into AI solution designs, model choices and implementation approaches. Combine architecture with hands-on engineering, production analysis and service improvements.
Solution architecture
Translate business use cases into solution designs across enterprise data, machine learning, LLM platforms and agentic workflows.
Model selection and implementation
Guide model selection and implementation choices, balancing functional requirements, cost, capacity and operational constraints.
Shared AI service controls
Contribute to the design and implementation of shared AI service controls, including usage budgets, quotas, rate limits and tool-calling integrations.
View 3 more contributions
Production reliability and scaling
Investigate production reliability and scaling issues, contributing code changes and service enhancements alongside platform engineering teams.
AI service adoption
Help client teams adopt AI services through architectural guidance, implementation support and production troubleshooting.
Operational support
Participate in a shared operational support rotation across enterprise AI platforms.
Personal engineering kept separate from client work: an LLM reference architecture, Ondine and execution-control tools for agents. This lets me publish code, experiments and their limits without exposing a client system.
Vauban, reference architecture and evidence lab
A vendor-neutral architecture for testing gateways, guardrails, tool policy, approval gates and trajectory logging. Public claims point to a lab artifact or experiment.
Controlled measurement instead of vendor claims
Across 680 runs and six attack classes, text filtering caught 27% of attacks versus 97% for a deterministic policy on tool arguments. The protocol and its limits are published.
ToolGuard, execution governance
A fail-closed sidecar that authorises agent tool arguments before execution and records each decision in a hash-chained, append-only ledger. 465 automated tests.
View 1 more contribution
Qadar, forensic engine for agent trajectories
Information-flow taint tracking and past-time LTL ordering contracts over tool-call trajectories, producing value-free cryptographic proofs on a tamper-evident ledger. 180+ tests. ToolGuard blocks; Qadar proves.
Built an AI procurement optimization platform deployed across more than seven countries. I contributed critical code and led architecture choices across recommendation, LLM validation and production engineering.
LLM Product Validation Service
Architected and shipped a distributed LLM verification service built from multiple Ondine instances, with Azure OpenAI accessed through LiteLLM. The worker fleet processed DataFrame rows asynchronously and scaled to 2,000 concurrent checks across the deployment. It automated the review of AI-generated product-swap pairs for the Italian market, leaving humans to handle exceptions.
Agentic Procurement: Closing the Savings Loop
Designed an agentic loop connecting recommendations to execution: propose compliant substitutes from live catalog data, explain off-contract purchases and rank fallback options when a supplier runs out of stock.
Hard Safety Constraints: Allergens as Fail-Closed Rules
A wrong food substitution is not a bad quarter, it is a health incident. Allergen and nutrition specifications are enforced as deterministic fail-closed checks outside the model, never as prompt instructions, and no proposal reaches the ERP without human approval and an audit record of who approved what. Bank-grade agent controls applied in an industry nobody calls regulated.
View 5 more contributions
BERT Embedding Pipelines & Graph Database
Maintained and optimized BERT embedding pipelines (ONNX Runtime) combined with MCA-based categorical encoding for product similarity computation at scale. Results stored in ArangoDB (graph database) enabling real-time product recommendation queries via a FastAPI swap API.
Reusable Multi-Region Data Assets
Defined and engineered data preprocessing pipelines as reusable assets with a shared core and market-specific extensions (North America, UK & Ireland, Italy, Brazil, Australia), each with distinct business rules, data sources (SAP, Dremio, SpendIQ), supplier aggregation logic, and geolocation enrichment. Modular architecture: build once, adapt per country. Optimized memory footprint and Databricks compute costs.
Ondine Integration Design
Designed the integration of Ondine (an open-source LLM batch processing framework) to replace legacy regex-based product attribute extraction, targeting an improvement from 60–70% to 90–95% precision on structured attribute parsing from unstructured product descriptions.
CI/CD, Observability & MLOps
Led the full Pydantic v2 migration across the codebase including custom Databricks Docker images for CI/CD compatibility. Built SdxMonitor, the production-grade ML monitoring infrastructure: embedding drift detection (cosine similarity distribution shifts), swap acceptance rate tracking per market, SLA alerts on pipeline latency, model versioning with MLflow, data lineage tracking across pipeline stages, and compute cost dashboards aggregated across 7 markets. Collaborated with SRE practices for operability, reliability, and cost control integration. Automated test gates in Azure DevOps for model performance regression.
LLM-Powered Product Knowledge RAG for Swap Enrichment
Deployed an on-premise GPU cluster running open-source LLMs (DeepSeek-R1-Distill, Qwen-1.5B) to build a RAG system over product facet knowledge bases, extracting attributes (allergens, nutritional profiles, packaging specs, certifications) that BERT similarity alone couldn't capture. Hybrid retrieval (Elasticsearch BM25 + Weaviate vectors + cross-encoder reranking) orchestrated via LangChain, with RAG quality evaluated through RAGAS metrics (faithfulness, context relevance, answer correctness). Fed enriched facet data into the Databricks ingestion pipelines (Brazil, Italy, North America) to improve swap pair quality before scoring.
Led the data architecture design for France's largest electricity producer and supplier, managing 35 million smart meters sending readings every 15 minutes, totaling 1.2 trillion readings per year.
Island Grid Datalake Architecture
Designed the complete datalake architecture on Cloudera CDH for insular electrical systems (French overseas territories). Data Vault-based design to progressively integrate new sources. Critical because island grids are isolated (no continental interconnection), so data reliability directly impacts grid stability.
Consumption Forecasting & Grid Balancing
Built the data layer feeding high-performance consumption forecasting models (GAM). The datalake supplies historical consumption, weather, and production data for real-time supply-demand balancing, especially critical with intermittent renewable energy sources.
Smart Meter Data Pipelines
Built high-throughput data pipelines processing readings from 35 million smart meters every 15 minutes. Apache NiFi for ingestion, HBase for real-time access, Hive for batch analytics, feeding consumption forecasting models and grid operations dashboards.
View 1 more contribution
Grid Fault Detection
Integrated consumption data, sensor alerts, and weather data for real-time grid incident detection and localization. Built Elasticsearch and Kibana dashboards for operations teams to visualize network events in real time.
Owned the data engineering workstream for the IFRS17 regulatory transformation at a major European insurance group. IFRS17, the new international accounting standard for insurance contracts, requires massive actuarial computations: re-evaluation of millions of contracts, cash-flow projections, discounting, and cohort grouping. Drove the centralization engine and data backbone architecture. Shaped the AI platform and observability stack while coaching the team on data engineering best practices.
IFRS17 Centralization Engine
Owned and delivered the core IFRS17 Centralization tool, a SaaS platform for actuaries and accountants. Centralizes contract data, applies IFRS17 calculations (CSM, LRC, risk adjustment), and produces regulatory reports. Drove the migration of legacy SAS scripts to 50+ Spark production jobs processing millions of records daily.
Data Backbone Platform: Platform as Product
Designed and built a modular Java-based Data Backbone platform, treated as a reusable internal product deployed across business units. Layered architecture with dedicated modules for Data Ingestion, Data Storage, Data Processing, Data Management (flow control, audit, lineage, error handling, traceability), and Platform Management (identity, security). Each module designed as a reusable, independently deployable asset. Powered by Kafka + MapR for streaming and Spark for batch processing.
AI Platform for Data Scientists
Helped build an internal AI platform enabling data scientists to spin up notebooks and workspaces on demand. Provided self-service compute environments with pre-configured ML libraries, GPU access, and integrated data connectors, accelerating model development and experimentation cycles.
View 2 more contributions
SAS → Open Source Big Data Migration
Led the cultural and technical migration from 20 years of SAS to Spark. Defined training programs, documentation standards, and ETL pipelines that reproduced SAS results to build trust with actuarial teams. Achieved 40% improvement in data quality and significant reduction in reporting latency.
Observability, SLAs & MLOps
Established the production-grade observability stack: Grafana dashboards for proactive pipeline monitoring with SLA alerts, Oozie orchestration for batch jobs, and Deequ-based automated quality checks (data contracts) to guarantee data integrity at each pipeline stage. Set standards for cost control and operability across the team.
Data and ML engineering for an international payments platform processing 5.8 million SEPA operations daily and serving more than 1,000 financial institutions.
Sanctions & KYC/AML Compliance Pipeline
Built data pipelines that ingest international sanctions lists (OFAC, EU, UN), cross-reference them against SWIFT messages in near real-time, and flag alerts. Millions of transactions per day filtered against hundreds of thousands of sanctions entries. Integrated the SWIFT KYC Registry and Sanctions List Distribution Service.
Fraud Detection on International Payments
Developed Spark pipelines for anomaly detection on correspondent banking flows, detecting unusual patterns such as a country sending 10× normal volume to a correspondent, or an IBAN appearing across multiple suspect corridors. Fed scoring models for risk assessment.
SWIFT GPI: End-to-End Payment Tracking
Contributed to the data lake centralizing SWIFT GPI tracking data for 11 countries and 200 payment corridors. Built real-time dashboards for operations teams and clients to monitor payment status, processing time, and fees at each stage.
View 2 more contributions
Data Anonymization Framework
Developed a Scala/Spark anonymization framework that masks sensitive payment data (IBAN, names, amounts) while preserving statistical properties for test environments. Results stored in Oracle for non-production use.
Paylib: Targeted Marketing & Appétence Prediction
Built a targeted marketing system for Paylib (the French mobile payment solution). A machine learning pipeline predicting user appétence to identify which customers would accept promotional offers, enabling precision targeting instead of mass communication and significantly improving conversion rates.