Writing

Faustus agent

Faustus is a CLI-first agentic data harness that unifies data engineering, quality, analytics, exploration, and ML workflows to take data from ingestion to reliable, production-ready outcomes.

People who work with data live in a paradox: our tools keep getting more powerful, but the end-to-end workflow is still fragmented across scripts, notebooks, jobs, manual validations, and scattered context.

That is why I started building Faustus: a CLI-first, agentic harness designed specifically for data professionals.

While many current harnesses focus on generic coding or personal productivity, the proposal here is different: tackle the full data lifecycle, with a focus on engineering, quality, analytics, exploration, and production ML.

The problem I want to solve

Today, even mature teams face similar pain points:

  • fragile pipelines that are hard to observe
  • validations that become manual tasks
  • prompts and automations without standards
  • rework caused by missing shared context
  • token costs growing without governance

In practice, this leads to delays, silent bugs, and less reliable decisions.

The Faustus vision

My goal is to build a real productivity layer for data, based on:

  • tools: practical utilities for recurring tasks
  • MCPs: structured integration with sources, services, and context
  • skills: reusable best practices for consistent execution

All with two core concerns: technical quality and cost efficiency.

Project principles

  • standards before improvisation
  • observability from day one
  • automation with guardrails
  • living documentation
  • token optimization without sacrificing outcomes

What comes first

In the next iterations, I will focus on:

  1. CLI foundation and the initial harness architecture
  2. first set of skills for data engineering and data quality
  3. minimal validation and troubleshooting flow
  4. usage and cost measurements per execution

Building in public

This post marks the beginning.

The idea is to evolve Faustus in public, sharing decisions, experiments, mistakes, and learnings along the way. If you are also dealing with data challenges in production, let us exchange ideas.