Skip to content

Repository files navigation

spendlabel

Benchmark: CPV category classification across multiple AI paradigms on real UK public procurement data, consumed from Confluent Kafka (cpv-raw) by per-paradigm Python consumers.

What It Does

SpendLabel takes real UK government contract notices (from Contracts Finder) and classifies each into one of 15 CPV 2-digit categories using 7 fundamentally different AI paradigms — then measures accuracy, latency, and throughput head-to-head on the same data stream.

Architecture

CSV files ──► Producer ──► Kafka (cpv-raw) ──► per-paradigm Kafka consumers ──► PostgreSQL ──► ui-chart dashboard
                                                │
                                                ├── hardcoded     (keyword/regex rules)
                                                ├── spark_ml      (TF-IDF + classifier)
                                                ├── deeplearning_onnx (neural net via ONNX)
                                                ├── solver        (OR-Tools constraints)
                                                ├── langchain     (LLM agent, zero-shot)
                                                ├── n8n           (webhook workflow)
                                                └── mcp           (Claude + tools)

Each consumer is a separate confluent-kafka consumer subscribed to cpv-raw, writing predictions to PostgreSQL via the shared consumer loop. Ground truth (CPV 2-digit prefix) is already in the dataset — no external labelling needed.

Tech Stack

Layer Technology
Consumers Confluent Kafka (confluent-kafka Python client)
Message broker Confluent Kafka (KRaft, dockerised)
ML PySpark, ONNX Runtime
Solver Google OR-Tools
LLM LangChain + OpenAI
Workflow n8n
Agentic MCP + Claude
Storage PostgreSQL
Dashboard Streamlit + Plotly
Data pandas, PyArrow

Project Structure

spendlabel/
├── data/raw/                       ← 3 CSV files (not committed)
├── services/
│   ├── producer/app/               ← CSV → Kafka publisher (local script, not dockerised)
│   │   └── publish.py
│   ├── kafka-consumers/app/       ← Service 1: Kafka consumers → PostgreSQL (dockerised)
│   │   ├── main.py                 ← entrypoint: python main.py --paradigm <p>
│   │   ├── config.py               ← Kafka settings (env-var overrides)
│   │   ├── consumers/
│   │   │   ├── consumer_loop.py    ← shared consume→classify→Postgres loop
│   │   │   ├── paradigms.py        ← single-source-of-truth paradigm registry
│   │   │   ├── cpv_labels.py       ← canonical CPV division label set
│   │   │   ├── hardcoded/          ← Rule-based classifier
│   │   │   ├── spark_ml/           ← Spark ML classifier
│   │   │   ├── deeplearning_onnx/      ← ONNX runtime classifier
│   │   │   ├── solver/             ← Constraint solver classifier
│   │   │   ├── langchain/          ← LangChain LLM classifier
│   │   │   ├── n8n/                ← n8n webhook classifier
│   │   │   └── mcp/                ← MCP + Claude classifier
│   │   ├── db/
│   │   │   ├── schema.sql          ← classifications + metrics tables
│   │   │   └── connection.py       ← psycopg2 pool + insert / materialise
│   │   ├── requirements.txt
│   │   └── Dockerfile
│   └── ui-chart/app/               ← Service 2: PostgreSQL → charts (Streamlit, dockerised)
│       ├── main.py
│       ├── charts/
│       │   ├── sunburst.py         ← CPV spend sunburst (off-contract colour)
│       │   └── accuracy.py         ← accuracy-per-paradigm table
│       ├── requirements.txt
│       └── Dockerfile
├── config/
│   └── confluent.env.example       ← Kafka + Postgres credentials template
├── docker-compose.yml              ← postgres + consumers + ui-chart
├── notebooks/                      ← EDA on the dataset
├── ATTRIBUTION.md                  ← Data source + licence
├── requirements.txt                ← producer + dev tooling
└── README.md

Quick Start

# 1. Create virtualenv
python -m venv .venv && source .venv/bin/activate

# 2. Install dependencies
pip install -r requirements.txt

# 3. Configure Kafka credentials
cp config/confluent.env.example config/.env
# Edit config/.env with your Confluent Cloud credentials

# 4. Place CSV files in data/raw/

# 5. Publish data (local script — not containerised)
cd services/producer/app && python publish.py && cd -

# 6. Bring up Postgres + consumers + dashboard
docker compose up --build

# Run a different paradigm:
#   docker compose run --rm consumers python main.py --paradigm spark_ml

# Dashboard: http://localhost:8501

Data

UK Contracts Finder notices via data.gov.uk — Open Government Licence v3.0. See ATTRIBUTION.md.

About

7 AI paradigms — hardcoded rules, Spark ML, Kafka Streams + ONNX, constraint solver, LangChain, n8n, MCP — run against the same CPV classification task on real UK procurement data. Same microservice, different brain. Accuracy, latency, cost, setup time compared. No winner fits all — but each has a clear home.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages