Hamza Ben Marzouk

Hamza Ben Marzouk

Data & AI Engineer · Paris

Français

About

Data & AI Engineer with 7+ years of experience, including 5 in consulting (OCTO Technology), across demanding industries: banking, insurance, energy, media and cloud. I design and harden data platforms (Snowflake, dbt, Databricks, Spark) and put AI into production: agents, LLM evaluation, adoption by business teams. I bridge business and engineering, upskill the teams I work with, and leave behind systems that are documented, tested and maintainable.

What I do

Experience

Wakam

Feb 2024 – Present

Data Platform Engineer

Back on the Data Platform to industrialise the platform's security, governance and reliability.

  • Designed a GitOps system for temporary Snowflake access: request by pull request, approval, automatic revocation and audit log, including for sensitive (GDPR) data
  • Moved Snowflake configuration to infrastructure as code (Terraform): network policies, warehouses, service accounts, security tasks; addressed CIS benchmark recommendations
  • Built a regulatory data mirror (UK data residency) to Azure Blob Storage, with an independent daily integrity check
  • Hardened Prefect orchestration: dev/prod isolation, late-run monitoring, Slack alerting
  • Migrated CI/CD from Azure DevOps to GitHub Actions, with ephemeral review environments per pull request, a SonarCloud quality gate and a Python 3.13 / uv migration
  • Agentic tooling for the team: Claude Code skills automating PRs, tickets, diagnostics and routine operations

Stack : Snowflake · Terraform · Prefect · dbt · Azure · GitHub Actions · Python · Claude Code

AI Engineer — KamAI team

Built and deployed AI solutions with a dual goal: industrialise AI agent evaluation and maximise adoption of AI tools across the company.

  • Designed and deployed an end-to-end evaluation system for business-critical AI agents
  • Automated generation of synthetic evaluation datasets, validated by business experts in a Retool app
  • Automated evaluation pipelines with centralised result tracking in Langfuse
  • Shipped a conversational chatbot on wakam.com for policyholders (agentic workflows on Dify)
  • Drove adoption of the Dust platform; ran bi-monthly hackathons (average satisfaction ≥ 4/5) and coached business teams on high-value use cases

Stack : Dust · Dify · Langfuse · Retool · LLMs · Python

Data Engineer — Data Platform

Two strategic workstreams, PDX (Partner Data eXchange) and DPF (Data Platform Foundation), ensuring the data platform's reliability, scalability and operability.

  • Technical lead on data contracts: brought 3 partner teams to autonomy in defining and applying the standards
  • Delivered a hardened end-to-end pipeline (partner exchange → Snowflake), with documentation, knowledge transfer and formalised processes
  • Optimised the monthly run: 50% of data support requests handled self-service
  • Full observability with Datadog (alerting, dashboards); rebuilt the data quality framework and introduced Elementary
  • Designed ETL / reverse ETL pipelines and restructured the dbt codebase by domain; significantly reduced data incidents

Stack : Snowflake · dbt · dlt · Databricks · Elementary · Datadog · Python

OCTO Technology

Feb 2019 – Jan 2024

Consultant — client engagements

Scaleway — Data Ops Consultant

Launch of Scaleway's first data product: a managed Spark offering (Spark as a Service).

  • Validated technical prerequisites: scalability, connectivity, resilience
  • Implemented and validated the first use cases (proof of concept)

Stack : Apache Spark · Spark as a Service · POC

Mobilize Financial Services (Renault Group) — Backend Developer

Creation of the Mobilize Pay neobank, alongside Accenture: built and shipped the mobile banking app.

  • Implemented banking features (layered architecture inspired by clean architecture)
  • Upheld engineering best practices and drove the team's continuous improvement

Stack : GCP · Java · Spring Boot · Flutter · PostgreSQL · GitLab CI

RelevanC — Data Engineer

Data lake overhaul (optimisation and remediation).

  • PySpark data cleansing and processing pipelines; CI/CD and infrastructure as code
  • Upskilled the team through pair and mob programming and code reviews

Stack : GCP · BigQuery · Apache Spark · Airflow · Dataproc · Terraform

SACEM — Tech Lead

Automated reconciliation of music rights contract updates submitted by publishers.

  • Designed ingestion, enrichment and business-rule workflows, plus process tracking metrics
  • Technical leadership: team facilitation, pair programming, code reviews, architecture design with the architects

Stack : AWS · Java · Python · PostgreSQL · Lambda · ECS · Terraform

Engie Digital — Backend Developer

Livin' smart city platform: real-time air quality, street lighting, traffic.

  • Designed and built product features, bringing data expertise on ingestion and processing
  • Drove DevOps culture and Accelerate practices; ran event storming workshops

Stack : AWS · Java · Spring Boot · PostgreSQL · MongoDB · Apache Spark · CQRS

BNP Paribas BDDF — Data Engineer

Rebuilt client file processing pipelines: migration from a DB2 mainframe to a Big Data stack.

  • Spark aggregation and transformation jobs, orchestrated with Oozie
  • Set up and maintained the CI/CD pipeline

Stack : Scala · Apache Spark · HDFS · Hive · Oozie · Jenkins

Advisory engagements

Skills

Data engineering
Snowflake dbt dlt Databricks Apache Spark Delta Lake Kafka BigQuery Airflow Prefect
Data quality & observability
Elementary Datadog Data contracts Alerting
AI & LLMs
AI agents LLM evaluation Synthetic datasets Dust Dify Langfuse Retool Claude Code
Cloud & infrastructure
AWS Azure GCP Terraform GitHub Actions GitLab CI Docker
Programming
Python SQL Java Scala
Architecture & practices
Data Mesh Data modelling Clean Architecture Software Craftsmanship Accelerate Agile

Education

  • Master's degree, Artificial Intelligence, Systems & Data
    Université Paris Dauphine · Sept 2017 – Sept 2018
  • Engineering degree, Computer networks & telecommunications
    INSAT, Tunis · Sept 2012 – Sept 2017

Certifications

  • Databricks Certified Data Engineer Associate (verify)
  • Databricks Certified Associate Developer for Apache Spark 3.0 (verify)
  • Databricks Partner Training — Solutions Architect Essentials (verify)
  • AWS Certified Solutions Architect — Associate

Languages

  • French — Bilingual
  • English — Professional

R&D & talks

  • Frugal architectures · Framework for assessing software architectures through their carbon footprint, with actionable recommendations.
  • CPU cache · Tech talk: cache lines, spatial and temporal locality, cache coherence.
  • Databricks platform · In-depth exploration of the Databricks ecosystem and internal knowledge sharing.

Interests

Contact

The best way to reach me is by email. [email protected]