Ivan
Shamaev
I design DWH and DataLake, build ETL/ELT pipelines — and automate the data team's work with AI agents (RAG, text-to-SQL, harness). 6+ years in data warehouses and platforms, 15+ years in data — from BI consulting to DWH architecture in e-commerce.
Data Engineer + AI agents for data engineering
Senior Data Engineer / AI Data Engineer. Currently at Ozon Tech: batch pipelines on Vertica, Trino and PySpark, DWH domains on the anchor model, and AI agents that automate data engineering.
My path started in financial consulting (SAS, Oracle Hyperion) and Qlik development, went through a 3-layer DataLake and BI platforms on ClickHouse, Qlik Sense and Apache Superset, and today it is focused on the modern data stack — dbt, Airflow, Trino, ClickHouse, Vertica — and the practical use of LLMs in data engineering.
Data Engineering
Batch ETL/ELT pipelines, DWH architecture (Anchor Modeling), data marts, data quality, migrations and optimization.
AI Agents
RAG pipelines, LLM integration, knowledge-base agents, SQL copilots, metadata-driven automation.
Technology stack
Data engineering, DWH, AI and infrastructure tools — ordered by market demand
15+ years in data
From financial consulting and BI platforms to DWH architecture, DataLake and AI agents
— present
- Build the DWH data platform: business flows in the DDS layer on the anchor model (Anchor Modeling), batch ETL pipelines on Airflow (Vertica, Trino, PySpark) — 40+ key data marts / 20+ domains
- Optimize complex analytical SQL queries on Vertica — up to 30–50% faster: rework join patterns and eliminate static planner defects
- Migrate heavy ETL processes Vertica → Trino: queries that failed with out-of-memory (50 GB/node) now run reliably after join optimization and Trino session-parameter tuning
- Build Airflow DAGs (dag-factory) (Vertica, Trino, PySpark, HDFS; incremental-load templates) and MSSQL ↔ Vertica integrations; ad-hoc extraction of large volumes on PySpark + Hadoop/HDFS (MPP / distributed computing)
- Run Data Quality ETL monitoring in Grafana with control of key failure points across the DWH domain in MSSQL
- As project tech lead run code review and design review; on on-call duty — ETL support (sensors, restarts, data arrival) and user support, code review, disaster-recovery (DR) drills for data-center outages
- Design the architecture of a DWH AI agent MVP (RAG over metadata and table relations + text-to-SQL) as the DWH-side expert — access, performance, process R&D — so 100+ business users of the platform can self-serve answers and ready SQL, offloading data engineers
- Introduce agentic development (OpenCode, Claude Code, Cursor, MCP, specs/rules) and coach the team on tools and technologies
— Nov 2024
- Built and automated ELT pipelines for DWH data marts (MSSQL, ClickHouse): orchestration in Airflow, transformations in dbt (ELT) and Python — core layer and data marts (~20 marts)
- Led the BigQuery → Yandex Cloud migration (ClickHouse, Airflow, dbt) with zero downtime on key data marts; trained cross-functional teams on the new stack
- As team lead ran a team of 11 people (5 data engineers, 3 DWH analysts, 3 analysts); hired 4 (3 analysts + a BI team lead) and stabilized the DWH after data engineers left — restored Airflow loads, optimized MSSQL/ClickHouse, onboarded new hires
- Ran a storage and BI audit and optimization: moved tables to columnar format, removed views, formalized the refresh policy — fixed DataLens performance issues
- Set up MSSQL and DWH-infrastructure monitoring in Grafana; controlled data refresh in MSSQL / ClickHouse via Airflow
- Built data marts for e-commerce operational analytics and a metrics tree (semantic layer / OKR) as a single source for C-level reporting
— Nov 2023
- Designed and built a 3-layer Data Lake (raw → staging → marts): ETL from the Facebook Graph API in Python with daily refresh — ~9 sources / 0.6 TB of data
- Developed and maintained ETL for internal corporate systems (MySQL, Microsoft Navision, PostgreSQL)
- Modeled and built ClickHouse data marts (Data Lake marts layer) for financial and performance analytics (PnL, Balance Statement); plan-vs-actual tools cut reporting-prep time by −90%
- Introduced Apache Superset as a Qlik alternative — cut license costs and widened data access without budget growth
- Introduced a data catalog (OpenMetadata): cataloged Data Lake layer metadata and built a data lineage prototype — gave analysts transparency into data provenance (data governance)
- Developed custom Superset plugins (React/TypeScript) and set up CI/CD (GitLab CI) for automated Docker image builds; ran Superset migrations 1.3.2 → 2.1.1
- Developed Python SSE plugins for Qlik Sense (server-side extensions)
— Jul 2020
- Built ETL pipelines and data integrations from 1C ERP (accounting) and APIs (Bitrix24 CRM, Yandex Metrica, Google Analytics, Mango Office) — first in PHP, then rewritten in Python
- Developed sales-funnel analytics for the online storefront, intra-group PnL calculations, and a marketing-campaign evaluation app — counterparty activity grew several-fold
- Introduced code versioning in Git (.qvs architecture) and PowerShell automation
- Built a C# Windows Service to integrate QVS with the NPrinting API
- Award "Best Employee, Q4 2019"
— May 2017
- Implemented a budgeting system on Hyperion Planning; deployed QlikView EPM (analytics over the GOLD ERP)
- Developed data integrations and calculations in Hyperion cubes
- Optimized data loads and provided post-project support
— Nov 2014
- Supported and developed Oracle Hyperion Planning (calculations, integrations)
- Developed analytical data models in QlikView
— Aug 2013
- SAS Base, SAS FM, SAS ABM; client consulting, pre-sales; Oracle EPM, SAP BO PCM
— Sep 2011
- Supported the Galaktika system, designed business processes, wrote specifications for developers
Selected case studies
Results with measurable business impact
Built a Harness AI Agent Support Assistant on top of the DWH to answer questions from the data platform's business users:
- VectorDB Qdrant: Confluence data split into chunks and loaded into the vector store
- MCP: Confluence, Jira, DataHub for fetching up-to-date metadata, projects and tools
- Tooling: search across DWH repos (DAGs, tables, logical models, descriptive attributes)
Piloted a data-pipeline development approach on OpenCode and Claude Code using spec-driven development.
Led a full DWH migration to the cloud, team of 8. Rebuilt ETL pipelines and data marts on ClickHouse, Airflow and dbt; trained cross-functional teams on the new stack.
Result: product analytics moved to the new stack (DDS and DM layers migrated — 15+ data marts).
Plan-vs-actual analysis tools cut the time to prepare per-department performance reporting by 90%.
Selected and launched Superset as a complement to Qlik. Custom React/TypeScript plugins. GitLab CI image auto-builds. Migrations 1.3.2 → 2.1.1.
Grafana dashboards for monitoring critical failure points and pipeline health. 100% DWH-domain coverage.
Designed and built a 3-layer DataLake (raw → staging → marts): ETL from the Facebook Graph API on Python and s3, with daily refresh and ClickHouse data marts. Foundation for content-production operational analytics.
Education & growth
Let's
get in touch
If you have questions about my experience or an interesting project — reach out any way that works for you.