Tuesday, August 18, 2026

From DBA or Data Engineer to AI Engineer: A Realistic Path

If you spend your days tuning queries, managing pipelines, or keeping a production database alive, you already carry most of what an AI engineering role needs. What is missing is not a new career from scratch, it is a specific set of additions on top of what you already do well. Here is the path that actually gets you there, without pretending you need to become a research scientist first.

1. Treat Python as your primary language, not a scripting add-on

Most DBAs and data engineers already write Python for ETL glue code, but AI engineering asks for more: comfort with async code, working with SDKs like OpenAI's or Anthropic's, and writing code that calls external APIs reliably, with retries and error handling. If your Python has mostly lived inside pandas scripts, push it into building small services and command line tools before touching anything AI-specific.

2. Learn how LLMs actually work, at the level you need to use them well

You do not need to understand transformer internals to be effective. You do need to understand tokens and context windows, the difference between a base model and an instruction-tuned one, prompt structure, temperature and sampling parameters, and function or tool calling. This is the layer that separates someone who can call an API from someone who can design a reliable system around one.

3. Build on the database skills you already have: vector search and RAG

This is where your background gives you a real head start. Retrieval Augmented Generation is fundamentally a data engineering problem with a language model bolted on the end: chunk documents, generate embeddings, store and query vectors, and feed the results into a prompt. If you already know PostgreSQL, pgvector gets you there without learning a new database engine. If you work in Redshift or SQL Server environments, understanding how vector search differs from traditional indexing will make you far more useful on an AI team than someone coming in without any database grounding.

4. Get hands-on with a managed AI platform on the cloud you already use

If your infrastructure background is AWS, go deep on Amazon Bedrock: model access, knowledge bases, and agents. If you live in Azure, do the same with Azure OpenAI Service. The point is not to learn every platform, it is to learn one well enough to provision it, secure it with proper IAM, and monitor its cost and usage the way you already monitor a database.

5. Ship one project that proves you can build, not just describe

Interviews for AI engineering roles increasingly ask for a working example over a certificate. A small RAG chatbot over your own documents, using a vector-enabled Postgres and a Bedrock or Azure OpenAI model, is enough to demonstrate the full pipeline: ingestion, embedding, retrieval, and generation. I laid out this exact project, along with four others that mix cloud data warehousing and AI, in an earlier post on starter projects for an AI and data engineering portfolio.

6. Learn to evaluate and monitor AI systems, not just build them

A model that answers well in a demo can fail quietly in production: hallucinated answers, drifting retrieval quality, or cost spikes from runaway token usage. Employers increasingly want people who can set up evaluation pipelines and observability around an AI system, not just wire the initial integration. This is a natural extension of the monitoring instincts a DBA or data engineer already has, applied to a new kind of workload.

What This Path Actually Buys You

You are not competing with machine learning researchers for these roles, and you do not need to be one. Most AI engineering work in production is systems work: reliable pipelines, sound data models, and careful integration of a model into something people use. That is the job you have been doing all along, with a new component added on top. Start with the RAG project, since it touches every skill above at once, and let the rest follow from there.

If you are setting up a machine specifically for this kind of work, I covered the full setup, from picking a model to getting PostgreSQL, Docker, and PyTorch running natively, in an earlier post on configuring a Mac for data engineering and AI work.

Wednesday, August 12, 2026

A Guatemalan's Guide to Drinking Coffee Properly

My mug
I grew up surrounded by coffee farms in Guatemala and somehow spent most of my adult life drinking whatever was closest, usually instant, usually while standing over the sink. If that sounds familiar, this is for you. No lectures, no gear worship, just why this coffee is actually worth the extra five minutes, how to pick a bag based on what you like instead of what a label tells you to like, and which machines are actually worth owning depending on how deep you want to go.


Why Guatemalan coffee earns the hype

Guatemala grows coffee on the slopes of active volcanoes, which is either the most metal origin story in the coffee world or just Tuesday for Guatemalan farmers. Volcanic soil is packed with minerals, the highlands stay cool even close to the equator, and slow growth at altitude gives the bean more time to build sugar and flavor before harvest. The regions genuinely taste different from each other too. Antigua leans chocolatey with a citrus edge, Huehuetenango turns bright and almost wine-like, Cobán goes softer and more floral thanks to all the rain. This is the actual, physical reason a bag from one region tastes nothing like a bag from another, not marketing copy dressed up to sound scientific.

Pick your coffee like you pick your music, based on taste, not trends

The single biggest mistake people make buying coffee is grabbing whatever bag has the nicest packaging. What actually matters is what you like in a cup, and that comes down to roast level.

If you like brightness, something with a little zing that wakes you up before the caffeine even kicks in, go lighter roast. Light roasts preserve more of the bean's natural acidity and fruitiness, they don't taste sour, they taste alive.

If you want something sweeter, rounder, and more familiar, medium roast is the safe, correct answer for most people, and it happens to be what most of the Huehuetenango bags below are roasted to. Nobody is judging you for liking the coffee equivalent of a comfortable couch.

Bags worth actually ordering

As I live in Guatemala, the brands usually available here are different from the ones on the international market, so i tried to find some of the brands that resemble the ones I prefer:

Fresh Roasted Coffee, Guatemala Huehuetenango, medium roast, single origin, kosher, sold in a 2-pack of 2 lb bags. Good option if you drink enough coffee that buying in bulk actually makes sense.

Guatemala Coffee, Huehuetenango, whole bean, medium roast, freshly roasted, 5 lb bag. Same region, bigger bag, for the household that goes through coffee like it's water.

Two Volcanoes Guatemalan Coffee, whole bean, 1 lb. Smaller bag, good if you want to try something new without committing to five pounds of a coffee you've never tasted.

The French way, easiest option...

A Veken French Press, 34oz, no plastic touching the coffee, thickened borosilicate glass, skips the paper filter entirely. That means more body, more oils, more of that heavier mouthfeel, which suits a medium roast Huehuetenango particularly well. Cleanup is slightly more annoying, that's the tradeoff, and it doubles as a cold brew pot if you're into that.

The Italian way, the most versatile one...


The Bialetti Moka Express, 12-cup is not technically espresso, don't let anyone tell you otherwise, but it's strong, concentrated, and genuinely delicious over ice or with steamed milk. The 12-cup size means you're not making this twice for a full household. It also makes a very satisfying gurgling noise right before it's ready, which is honestly half the appeal.

The Japanese way, V60 and actual patience

The Hario V60 Pour Over Starter Set, size 02 is the classic for a reason, cheap, simple, and it makes brighter beans genuinely sing. Pair it with the Hario Buono electric gooseneck kettle for actual control over your pour instead of just dumping water and hoping, and keep a stack of Hario V60 paper filters, size 02 on hand since running out mid-morning is its own special kind of tragedy. This method forces you to stand there for three minutes doing nothing but pouring water in circles, which is either meditative or mildly annoying depending on your morning.

Going full home barista


The Ninja Luxe Café Mini is the interesting middle ground, it does both espresso and drip in one machine, with a built-in burr grinder, a precision scale, and a manual steam wand, so you get real control without needing three separate appliances on your counter.

If you'd rather push one button and get a finished drink without thinking about grind size ever again, the Jura E4 in piano black is a fully automatic machine built for exactly that, bean to cup with none of the fuss. The Jura ENA 4 in Nordic white does the same job in a smaller footprint, good if counter space is the actual constraint rather than budget.

So which one do you actually need

None of them, technically. A bag of good Guatemalan coffee and a French press will outperform bad beans in a thousand dollar machine every single time. Buy the coffee that matches what you actually like drinking, and pick whichever method fits your morning, rushed, ritualistic, or somewhere in between. The volcano already did the hard part. You're just not allowed to ruin it on your end... Also, do not listen to the purist that says coffee should be black without sugar, is your money, you choose the way to drink it: sugar, cocoa, milk, cinnamon... go ahead and try new things!

This post contains affiliate links. As an Amazon Associate, studyyourdata.com earns from qualifying purchases at no extra cost to you.

Friday, August 7, 2026

Narrow Task Decomposition in AI Pipeline Agents

A nightly load into sales_fact fails. Ask an orchestrating agent to break down the diagnosis and it will usually stop at the first two subtasks that come to mind: check the logs, check the schema. That decomposition is narrow, and a fix built on it can miss the real cause. The problem is not the model, it is that the coordinator never questioned its own first list before delegating.

1. See what narrow decomposition misses

A first-pass breakdown of "why did sales_fact fail" often lands on log-analyzer and schema-inspector and stops there. Both are reasonable, and both are incomplete. Neither one checks whether the upstream extract delivered bad or late data, whether another job holds a lock on the same table, or which downstream dashboards are now stale because the load never completed. A fix based on two subtasks treats the visible half of the failure as the whole failure.

2. Add a self-critique step to the coordinator prompt

The fix is a coordinator prompt that forces a second pass before any delegation happens: generate an initial subtask list, ask what is missing, add subtasks to cover the gap, and only then hand work to subagents.

You are a coordinator diagnosing an ETL failure. When
decomposing the task:

1. Generate an initial list of subtasks.
2. Ask yourself: what systems, data sources, or failure
   modes are missing from that list?
3. Add subtasks to cover those gaps.
4. Only then begin delegating to subagents.

For an ETL failure specifically, consider:
- the failing job AND its upstream dependencies
- a structural cause (schema) AND a data cause (quality)
- this run AND whether it is a recurring pattern
- the technical root cause AND the downstream impact
  

Run the sales_fact example through that loop and the initial two-item list grows to four: log-analyzer and schema-inspector, plus a source-data-validator that checks the upstream extract for null spikes or row-count drops, and a downstream-impact-checker that lists which dashboards or scheduled jobs depend on the table. The self-critique step costs one extra generation and catches the two subtasks that a narrow first pass always skips.

3. Define each subtask as a subagent

With Claude's Agent SDK, each subtask becomes an AgentDefinition, with a description Claude matches against the task and a prompt scoped to that one job.

from claude_agent_sdk import AgentDefinition

log_analyzer = AgentDefinition(
    description="Diagnoses ETL failures from pipeline logs. "
                "Use when a nightly load job fails and the "
                "cause is unclear.",
    prompt="Find the failing step and report the exact "
           "error message and line number. Do not propose "
           "a fix.",
    tools=["Read", "Grep", "Bash"],
    model="haiku",
)

schema_inspector = AgentDefinition(
    description="Checks table structure for drift against "
                "the last known good schema.",
    prompt="Compare current columns, types, and "
           "constraints to the last successful load.",
    tools=["Read", "Bash"],
)

source_data_validator = AgentDefinition(
    description="Checks the upstream extract for the failed "
                "load. Use for row-count drops or null spikes.",
    prompt="Compare today's source row count and null "
           "rate to the 7-day average.",
    tools=["Read", "Bash"],
    model="haiku",
)

downstream_impact_checker = AgentDefinition(
    description="Lists dashboards and jobs that depend on "
                "the table that failed to load.",
    prompt="Find consumers of sales_fact and flag which "
           "are now stale.",
    tools=["Read", "Grep"],
)
  

AgentDefinition also takes a model field, so a subtask that is closer to pattern matching, like reading a log file or comparing two row counts, can run on a cheaper model while one that needs judgment, like assessing schema drift or downstream impact, stays on the default. That is the same idea covered in matching the model to the task, not the app, applied at the subagent level instead of a single orchestrator.

4. Wire the coordinator to call the agents

The coordinator prompt from step 2 becomes the system_prompt, and the four AgentDefinition objects from step 3 go into agents, keyed by the name Claude will use to call each one. Claude decides on its own which of the four to invoke, based on how the task matches each description.

from claude_agent_sdk import query, ClaudeAgentOptions

AGENTS = {
    "log-analyzer": log_analyzer,
    "schema-inspector": schema_inspector,
    "source-data-validator": source_data_validator,
    "downstream-impact-checker": downstream_impact_checker,
}

async for message in query(
    prompt="The sales_fact load failed last night, find out why.",
    options=ClaudeAgentOptions(
        system_prompt=COORDINATOR_PROMPT,
        # "Agent" auto-approves each subagent call
        allowed_tools=["Read", "Grep", "Bash", "Agent"],
        agents=AGENTS,
    ),
):
    for block in getattr(message, "content", None) or []:
        if getattr(block, "name", None) == "Agent":
            print("routed to:", block.input.get("subagent_type"))
  

Claude does not call a Python function directly, it emits a tool_use block named Agent, and block.input["subagent_type"] holds the name of the agent it picked, one of the four keys in AGENTS. Logging that value is how you confirm the coordinator actually routed to source-data-validator or downstream-impact-checker instead of quietly falling back to the two obvious subtasks. To force a specific agent instead of letting Claude choose, name it in the prompt: "use the downstream-impact-checker agent to list what depends on sales_fact."

5. The same routing in OpenAI's Agents SDK

OpenAI's Agents SDK expresses the same coordinator through handoffs. The triage agent's instructions carry the same four-step self-critique, and each of the four specialists is passed in its handoffs list, exposed to the model as a callable tool named transfer_to_<agent_name>. When Runner.run finishes, result.new_items contains a HandoffOutputItem for any handoff that fired, and its target_agent.name is the concrete signal, the OpenAI equivalent of reading subagent_type off Claude's Agent tool call.

6. One failure, traced end to end

Take one run instead of a list of what each agent can do. At 02:14, the nightly load into sales_fact fails, and the engineer on call prompts the coordinator with "the sales_fact load failed last night, find out why." The coordinator's first move is always the same: call log-analyzer, since nothing else can be scoped correctly before the actual error is known.

routed to: log-analyzer
  MERGE INTO sales_fact failed at 02:14:07 with:
  duplicate key value violates unique constraint
  "sales_fact_pkey" (order_id, sale_date)
  

A duplicate key on a merge has two plausible explanations: something changed about the table that now lets duplicates in, or the source data itself contains rows that collide on that key. Both are live possibilities under the self-critique step from earlier, so the coordinator calls schema-inspector and source-data-validator next. It does not call downstream-impact-checker yet, since nothing so far says anything about who reads this table.

routed to: schema-inspector
  no DDL changes on sales_fact in the last 14 days

routed to: source-data-validator
  source row count: 41,209 (7-day avg: 68,450, -40%)
  null rate on customer_id: 6.2% (7-day avg: 0.1%)
  

schema-inspector rules out a structural cause. source-data-validator explains the collision: a partial extract with a null spike on customer_id delivered rows that collapsed onto the same order_id and sale_date, tripping the unique constraint on merge. The coordinator stops here and never calls downstream-impact-checker, because the merge failed and rolled back, so nothing downstream ever read the bad batch. That agent earns its place in a run where a bad load succeeds silently instead of failing loudly, a different failure mode from this one.

Three of the four candidate agents ran, in two rounds instead of one batch, and the proposed fix follows directly from what they found: quarantine the extract that arrived overnight rather than retrying the load as is, and add a row-count and null-rate check ahead of the merge so a batch shaped like this one gets rejected before it reaches sales_fact, not after.

Before You Trust the Decomposition

Test the coordinator prompt against a case where you already know the full list of causes, and check whether its self-critique step actually surfaces the ones a naive first pass would miss. If it keeps landing on the same two or three subtasks regardless of the failure, the gap-checking questions in the prompt are too generic and need to name the specific systems in your pipeline.

Not every failure needs any of this. A connection timeout, a transient throttling error, a job that fails because another job it depends on is still running, these have deterministic fixes: retry with backoff, wait and requeue, alert and stop. Route the alert through a cheap classifier first, or a plain if-statement on the error code, and reserve the coordinator and its agents for failures that pattern-matching cannot already explain. Calling four agents to conclude "retry it" is not decomposition, it is waste.

Measuring RAG Solutions: Are We Retrieving the Right Information?

A RAG pipeline that answers questions in the demo is not the same thing as a RAG pipeline that answers them correctly.  The gap between the...