Skip to content
View jaimejrs's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report jaimejrs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaimejrs/README.md

LinkedIn Medium Gmail

gif1 gif2

About Me:

Data Engineering Analytics Engineering

"Be curious. Read widely. Try new things. I think a lot of what people call intelligence boils down to curiosity" - Aaron Swartz

I am a Data Engineer focused on Data Lakehouse architectures, data virtualization, and the Modern Data Stack. I specialize in building automated, scalable, high-performance pipelines, ingesting, transforming, and modeling data with an emphasis on quality, reproducibility, and reliable data quality checks. I am also a postgraduate student in Data Science and Artificial Intelligence.

Currently, I work as a Strategy and Performance Analyst at FIEC, acting as a data and process developer. I map business requirements into structured OKR/KPI models, prototyping deliverables in Figma and Excel before development. I build data extraction routines with Web Scraping, APIs, Python, SQL, and Power Query, unifying multiple sources into semantic models, and I develop Python scripts to automate Data Quality checks and statistical monitoring of report integrity. I also maintain Power BI semantic models with DAX measures for organization-wide dashboards.

Previously, as a Data Engineering Intern at Scanntech, I supported large-scale distributed processing of sell-out data with Azure Databricks, structuring the foundations of a Data Lakehouse architecture. I developed and monitored ELT flows with Apache Airflow, orchestrating data across lake layers, and used Dremio for data virtualization, building the semantic layer for direct SQL querying over the Data Lake. I also worked with Git for version control and CI/CD pipelines for safe, automated data pipeline deployments.

Alongside my corporate roles, I am a Professor at Digital College, where I teach "Data Analytics with AI" and "Advanced Python for Data Engineering." I guide students through practical projects involving the construction of Data Warehouses, pipeline orchestration, Streamlit applications, Machine Learning models, and AI agents. I also lead practical implementations of containerized data ecosystems with Docker, distributed processing in Hadoop, and task automation via Airflow DAGs.

📌 Featured Projects

Project Description Stack
data-engineering-lab-PY03 End-to-end lakehouse on Docker/Ubuntu Server: Airflow-orchestrated ingestion, Iceberg tables cataloged in Hive Metastore, federated queries via Trino, dbt modeling, and Superset dashboards. Airflow Spark Iceberg Hive Trino dbt
full_dataengineering Full ETL pipeline for sales analytics with geolocation enrichment, orchestrated with Airflow and backed by AWS Redshift and Hadoop/HDFS. Airflow AWS Redshift Hadoop Docker
movielens_project Analytics engineering pipeline that cleans, models, and serves the MovieLens dataset through BigQuery and Metabase dashboards. BigQuery Docker Metabase
engenharia_dados_meteorologicos Containerized ETL pipeline ingesting and transforming weather data for Fortaleza (CE), orchestrated with Airflow DAGs. Airflow Docker

To learn more about my projects visit my Data Portfolio

👨🏽‍🎓 Viral Analitycs / Data & ML 👨🏽‍💻 Data Analisys

Tools:

Orchestration
Apache Airflow Logo Apache Hop

Processing & Lakehouse
Apache Spark Apache Iceberg Hive Metastore Trino dbt Hadoop HDFS Databricks

Cloud & Storage
aws Logo Amazon S3 Redshift Google Cloud Storage Google BigQuery

Databases
PostgreSQL Logo MySQL Logo SQL Server Logo Oracle ODI

BI & Visualization
Power BI Logo Metabase Streamlit Logo

Languages & ML
Python Pandas Logo NumPy Logo Scikit-Learn Logo MLflow OpenAI Shell

Infra & Tooling
docker Logo Linux Ubuntu Server git Logo github Logo VS Code Logo Jupyter Logo colab Logo pacote office Logo

Management:


Pinned Loading

  1. tiktok_and_youtube-shorts_analysis tiktok_and_youtube-shorts_analysis Public

    Projeto para análise de dados de vídeos virais (TikTok e YT Shorts) estimulado pela cadeira de Business Analytics da Universidade Federal do Ceará.

    Python

  2. engenharia_dados_meteorologicos engenharia_dados_meteorologicos Public

    Pipeline completo de Extração, Transformação e Carga (ETL) de dados meteorológicos da cidade de Fortaleza (CE). O projeto foi construído para capturar dados climáticos em tempo real, tratá-los e ar…

    Jupyter Notebook 1

  3. movielens_project movielens_project Public

    Pipeline de Engenharia de Analytics end-to-end focado em processar, limpar e modelar dados do dataset MovieLens. O objetivo é extrair insights sobre popularidade de filmes, tendências de avaliações…

    Shell 1

  4. spanish-wine_datascience_analasys spanish-wine_datascience_analasys Public

    Projeto de análise e modelagem de dados relacionado a vinhos espanhóis. O objetivo deste projeto é explorar as características dos vinhos, extrair insights valiosos do conjunto de dados e avaliar m…

    Jupyter Notebook 1 2

  5. full_dataengineering full_dataengineering Public

    Pipeline ETL completo para análise de vendas com suporte a Pessoa Física e Jurídica, geolocalização por estado/cidade, orquestração via Airflow, Data Lake distribuído em HDFS e dashboard interativo.

    Jupyter Notebook

  6. data-engineering-lab-PY03 data-engineering-lab-PY03 Public

    Trabalho final do curso de Python Avançado para Engenharia de Dados e IA da Digital College

    Python