Automated notebook optimization for Databricks.
Finds your slowest notebooks, generates optimized versions with an LLM, validates them, and asks for your approval — all on autopilot.
Orchestrator job finishes
│
▼
┌──────────────────────────────────────────────────────────────┐
│ datadoc (runs as the last task in your job) │
│ │
│ 1. Jobs API ──▶ average of N runs ──▶ top K slowest │
│ 2. Export source + compute SHA-256 hash │
│ 3. Classify ──▶ 🟢 green / 🟡 yellow (tier + side-effects)│
│ 4. Estimate table sizes ──▶ broadcast safety map │
│ 5. Fetch last 5 rejected proposals for this notebook │
│ 6. LLM call ──▶ JSON with per-cell optimization diffs │
│ └─ rejected history injected as negative context │
│ 7. Reconstruct v2 by applying diffs to original source │
│ 8. Save v2 to {workspace_path}/proposals/ │
│ 9. If 🟢 green ──▶ validate output equivalence (T1/T2/T3) │
│ 10. Insert into datadoc.proposals (with source hash) │
│ 11. Notify team ──▶ Slack with approve/reject buttons │
└──────────────────────────────────────────────────────────────┘
│
▼ (when you decide)
┌──────────────────────────────────────────────────────────────┐
│ datadoc_approve │
│ ✅ approve ──▶ verify source hash ──▶ backup + apply v2 │
│ (if hash mismatch → reject as 'stale') │
│ ❌ reject ──▶ mark as rejected (feeds future LLM context) │
│ ↩️ rollback ──▶ restore from backup │
└──────────────────────────────────────────────────────────────┘
- Databricks workspace with an all-purpose cluster
- A Databricks serving endpoint for your LLM — recommended:
databricks-claude-opus-4-7 - A Slack app with
chat:writescope + bot token in a Databricks secret (or setnotifications.type: noneto skip Slack)
git clone https://github.com/404mqs/datadoctor
cd datadoctor
pip install pyyamlOpen datadoctor_config.yml and set:
databricks:
host: "https://<your-workspace>.azuredatabricks.net"
job_id: 123456789 # your orchestrator job ID
cluster_id: "xxxx-yyy" # all-purpose cluster
llm:
endpoint_name: "databricks-claude-opus-4-7"
notifications:
type: "slack"
slack:
channel_proposals: "C..."
channel_audit: "C..."
secret_scope: "datadoctor"
secret_key: "slack_bot_token"python scripts/upload_all.pyRun schema/datadoc.sql in a Databricks notebook or %sql cell.
python scripts/add_to_job.py # preview
python scripts/add_to_job.py --apply # applyThat's it. Data Doctor runs automatically at the end of your next orchestrator job.
When Data Doctor finds optimizations it posts to Slack:
| Button | Action |
|---|---|
| 🔍 Review v2 | Opens the proposed notebook in your workspace |
| ✅ Approve | Verifies the notebook wasn't edited since proposal; backs up + overwrites with v2 |
| ❌ Reject | Marks as rejected; description is fed back to LLM on the next proposal for this notebook |
Stale protection: if the notebook was manually edited after the proposal was generated, the approve action is blocked and the proposal is marked stale. Data Doctor will generate a fresh proposal the next day incorporating the manual changes.
To approve without Slack, run datadoc_approve manually and set the proposal_id widget.
To rollback after approving, run datadoc_approve with action=rollback.
Data Doctor is designed for continuous, incremental optimization — not a one-shot pass. Each run it targets the slowest notebooks that haven't been recently touched, applies one improvement, and moves on to the next bottleneck on the following run.
To support this, every task that receives an approved optimization enters a cooldown period (default: 5 days). During cooldown, that task is skipped from the slow-task ranking so Data Doctor can focus on other notebooks instead of repeatedly re-analyzing the same one.
This also prevents a subtle feedback loop: a freshly optimized notebook may run slower for the first few executions (Spark JIT warm-up, Delta cache cold start) and would otherwise keep appearing as a candidate — generating redundant proposals before the improvement has had time to prove itself.
To disable cooldown entirely, set agent.cooldown_days: 0 in the config.
| Tier | Meaning | What Data Doctor does |
|---|---|---|
| 🟢 Green | Writes Delta tables, no external side effects | Generates v2 + auto-validates + shows approve button |
| 🟡 Yellow | Has side effects (Sheets, Slack, APIs) or self-referencing tables | Generates v2 but skips auto-validation — requires manual review |
Built-in:
slack— Block Kit message with approve/reject buttonsnone— logs to stdout; proposals saved indatadoc.proposals
Custom notifier (Teams, email, etc.):
- Create
modules/notifiers/your_notifier.pyextendingBaseNotifier - Implement
send_proposal()andsend_result() - Set
notifications.type: your_notifierin the config
| Key | Default | Description |
|---|---|---|
databricks.host |
— | Workspace URL (required) |
databricks.job_id |
— | Orchestrator job ID (required) |
databricks.cluster_id |
— | All-purpose cluster ID (required) |
databricks.workspace_path |
/Workspace/Shared/DataDoctor |
Where files are uploaded |
agent.top_n |
3 |
Notebooks to analyze per run |
agent.cooldown_days |
5 |
Days to skip a task after applying an optimization |
agent.odd_days_only |
true |
Run every other day to save LLM tokens |
agent.delta_schema |
datadoc |
Delta schema for proposals and applied changes |
agent.self_task_key |
DATADOCTOR |
Task key to exclude from slow-task ranking |
agent.runs_to_average |
2 |
Past runs to average when ranking slow notebooks — higher values smooth out outliers |
llm.endpoint_name |
databricks-claude-opus-4-7 |
LLM serving endpoint |
llm.max_tokens |
16000 |
Max tokens for LLM response |
notifications.type |
slack |
slack or none |
notifications.slack.channel_proposals |
— | Channel for optimization proposals |
notifications.slack.channel_audit |
— | Channel for approve/reject confirmations |
notifications.slack.secret_scope |
datadoctor |
Databricks secret scope name |
notifications.slack.secret_key |
slack_bot_token |
Key within the scope |
cluster_cost_weights |
(absent) | Map of node_type_id → relative cost for ROI-based ranking (optional) |
photon_cost_multiplier |
1.5 |
DBU multiplier for Photon-enabled clusters |
By default, Data Doctor ranks notebooks by average duration. If your orchestrator uses clusters with very different cost profiles (a 16-worker memory-optimized cluster vs a single baseline node), duration alone may not reflect actual spend.
Add cluster_cost_weights to your config to rank by estimated DBU cost instead:
cluster_cost_weights:
Standard_DS3_v2: 1.0 # baseline
Standard_E8d_v4: 3.0 # memory-optimized, 8 cores
Standard_E16d_v4: 6.0 # memory-optimized, 16 cores
photon_cost_multiplier: 1.5 # clusters with Photon cost ~1.5× more DBUsScore formula: duration_min × node_weight × (num_workers + 1) × photon_mult
The ×3.0 multiplier appears next to the duration in Slack only when it differs
from 1.0 — no visual noise when all tasks share the same cluster type.
If cluster_cost_weights is absent from the config, ranking is unchanged.
Data Doctor supports multiple orchestrator jobs in the same workspace. Every proposal and query is scoped by job_id, so histories, cooldowns, rejection context, and performance gains are fully isolated per job — even when all instances share a single Delta schema.
Recommended: shared schema, separate configs — simplest to operate.
# config for the data pipeline job
databricks:
job_id: 111111111
agent:
delta_schema: "datadoc" # shared schema
self_task_key: "DATADOCTOR"
# config for the ML job
databricks:
job_id: 222222222
agent:
delta_schema: "datadoc" # same schema, isolated by job_id
self_task_key: "DATADOCTOR_ML"Alternative: separate schemas — maximum isolation, useful if you want per-team Delta permissions.
# ML job uses its own schema
agent:
delta_schema: "datadoc_ml"Run scripts/upload_all.py once per instance (it reads workspace_path and uploads to a separate folder per instance).
Data Doctor works with any OpenAI-compatible Databricks serving endpoint.
We recommend databricks-claude-opus-4-7 — it produces the most reliable diff-based JSON and handles complex PySpark notebooks well.
See webhook/README.md for setting up an Azure Function that enables true 1-click
approval from Slack — no Databricks UI required.
