Data Engineering Analytics Engineering
I am a Data Engineer focused on Data Lakehouse architectures, data virtualization, and the Modern Data Stack. I specialize in building automated, scalable, high-performance pipelines, ingesting, transforming, and modeling data with an emphasis on quality, reproducibility, and reliable data quality checks. I am also a postgraduate student in Data Science and Artificial Intelligence.
Currently, I work as a Strategy and Performance Analyst at FIEC, acting as a data and process developer. I map business requirements into structured OKR/KPI models, prototyping deliverables in Figma and Excel before development. I build data extraction routines with Web Scraping, APIs, Python, SQL, and Power Query, unifying multiple sources into semantic models, and I develop Python scripts to automate Data Quality checks and statistical monitoring of report integrity. I also maintain Power BI semantic models with DAX measures for organization-wide dashboards.
Previously, as a Data Engineering Intern at Scanntech, I supported large-scale distributed processing of sell-out data with Azure Databricks, structuring the foundations of a Data Lakehouse architecture. I developed and monitored ELT flows with Apache Airflow, orchestrating data across lake layers, and used Dremio for data virtualization, building the semantic layer for direct SQL querying over the Data Lake. I also worked with Git for version control and CI/CD pipelines for safe, automated data pipeline deployments.
Alongside my corporate roles, I am a Professor at Digital College, where I teach "Data Analytics with AI" and "Advanced Python for Data Engineering." I guide students through practical projects involving the construction of Data Warehouses, pipeline orchestration, Streamlit applications, Machine Learning models, and AI agents. I also lead practical implementations of containerized data ecosystems with Docker, distributed processing in Hadoop, and task automation via Airflow DAGs.
| Project | Description | Stack |
|---|---|---|
| data-engineering-lab-PY03 | End-to-end lakehouse on Docker/Ubuntu Server: Airflow-orchestrated ingestion, Iceberg tables cataloged in Hive Metastore, federated queries via Trino, dbt modeling, and Superset dashboards. | |
| full_dataengineering | Full ETL pipeline for sales analytics with geolocation enrichment, orchestrated with Airflow and backed by AWS Redshift and Hadoop/HDFS. | |
| movielens_project | Analytics engineering pipeline that cleans, models, and serves the MovieLens dataset through BigQuery and Metabase dashboards. | |
| engenharia_dados_meteorologicos | Containerized ETL pipeline ingesting and transforming weather data for Fortaleza (CE), orchestrated with Airflow DAGs. |
To learn more about my projects visit my Data Portfolio







