Skip to content
View Chiranjit-B's full-sized avatar
🎯
Actively Seeking Intern Data Intern Roles
🎯
Actively Seeking Intern Data Intern Roles

Block or report Chiranjit-B

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Chiranjit-B/README.md

Hi, I'm Chiranjit Banerjee 👋

Typing SVG

LinkedIn Gmail GitHub Profile Views


About Me

I started out as a data engineer at LTIMindtree, building PySpark pipelines and Snowflake data marts that kept enterprise reporting honest. Somewhere between fixing SCD Type 2 logic and chasing down bad joins, I got hooked on a bigger question: how do you make systems that don't just move data, but actually understand it?

That question is what pulled me toward AI. Today I split my time between cloud data engineering — Redshift warehouses, dbt models, Airflow DAGs — and applied AI engineering, building RAG pipelines, multi-agent orchestration layers, and LLM-powered products that ship to real users, not just notebooks.

Most recently, as a full-stack AI engineer, I've been architecting systems that stitch together voice AI, semantic search, and LLM extraction into single coherent pipelines — because the hard part was never the model, it's getting clean signal to it in the first place.

📍 Based in Boston, open to relocation | 📧 chiranjitstudies@gmail.com


⚡ Core Stack

AI & GenAI
LangChain RAG FAISS OpenAI Azure OpenAI
Cloud & Warehousing
AWS Azure Snowflake Databricks Redshift
Languages
Python SQL PySpark
Pipelines & Orchestration
Airflow Kafka dbt Docker
Analytics & BI
Power BI Tableau
Tools
FastAPI Streamlit Git

🚀 Featured Projects

🎬 Netflix Cloud Data Warehouse
Production-style ELT platform with medallion architecture, dimensional models, dbt tests, and full documentation, plus SCD Type 2 history tracking.
GCP Snowflake dbt SQL
⚖️ Legal Case RAG System
Semantic retrieval and RAG chatbot over 500+ legal documents, delivering context-aware query resolution through an interactive Streamlit app.
OpenAI FAISS Flask Streamlit Python
💳 Customer Retention Pipeline
Automated churn-analytics pipeline preparing warehouse-ready data and surfacing retention insights for the business.
AWS S3 Glue Athena Redshift Airflow Power BI
🏗️ Azure Customer Insights Platform
Scheduled Bronze–Silver–Gold pipeline moving customer and sales data from SQL Server into an analytics-ready cloud platform.
ADF ADLS Databricks Synapse Key Vault
📂 More Projects
Project Description Tech
Redfin Real Estate ETL Automation Cloud-native ETL pipeline ingesting 5M+ real estate records, cutting processing time by 70% via Airflow DAG optimization Python · Airflow · AWS S3 · Snowflake · Power BI
Real-Time Data Streaming Event-driven streaming pipeline for real-time data movement and processing Kafka · Python · Streaming

📊 GitHub Activity

Check out my contribution graph and pinned repos directly on my profile.

Let's build systems that turn raw data into decisions.

Pinned Loading

  1. Netflix-DataWarehousing Netflix-DataWarehousing Public

    End to End Data Warehousing Solution using GCP, Snowflake, DBT

  2. Customer-Retention-Insights-AWS-S3-Glue-Athena-Redshift-Power-BI Customer-Retention-Insights-AWS-S3-Glue-Athena-Redshift-Power-BI Public

    predicting customer retention using AWS cloud infrastructure and automating it using airflow

    Python 2

  3. Real-Estate-Data-ETL-Automation--AWS-Snowlflake-Airflow-PowerBI Real-Estate-Data-ETL-Automation--AWS-Snowlflake-Airflow-PowerBI Public

    This project develops an automated ETL pipeline for real estate data, using Apache Airflow for task scheduling, AWS S3 for data storage, and Snowflake for warehousing. It enables efficient data ing…

    Jupyter Notebook 1

  4. Motor-Vehicle-Collision-Analysis-Data-Warehousing-ETL-Automation Motor-Vehicle-Collision-Analysis-Data-Warehousing-ETL-Automation Public

    how advanced data engineering can transform public safety measures and enhance urban planning

  5. Azure-Data-Pipeline-for-Customer-Insights Azure-Data-Pipeline-for-Customer-Insights Public

    Developed a comprehensive data pipeline using Azure Data Factory, Databricks, and Synapse to extract, transform, and load customer and sales data from an on-premises SQL database into Azure. Implem…

    Jupyter Notebook 1

  6. Melatonin-Product-Sentiment-Analysis Melatonin-Product-Sentiment-Analysis Public

    Jupyter Notebook 2