Design · Measure · Decide — causal design, evaluation, and decision frameworks for enterprise AI

Designed & Measured AI

Posts

Turning AI and ML into decision making: governance, operating models, and essays on Designed & Measured AI. Newest first.

The business case for tabular foundation models

This essay extends questions I keep hearing in executive and alumni rooms: if build cost keeps falling, where does the value actually accrue? Tabular foundation models compress the same curve gen AI already did. The argument is that payoff sits in decision making and governance (which metric, who acts, what gets logged), not a smaller analytics budget. TabArena benchmarks, a taxonomy for where predictive signal lives, and readmission and trial forecasting composites.

What I'm paying attention to

A periodic scan across design, measurement, and decisions. Bookmarks from the week, not a forecast.

  • Design

    The lakehouse as agent operating layer. At Data + AI Summit 2026, Databricks framed the stack around context, control, choice, and cost: LTAP (single-copy transactional + analytical data), Unity AI Gateway for models/MCPs/agents with spend caps, and Agent Bricks as the developer surface. The bet is that agents read, loop, and write differently from human-facing software, so the data layer has to change with them.

    Parallel track: platform consolidation for production agents. The DataRobot Agent Workforce Platform pushes build, deploy, and govern into one lifecycle (identity-backed agents, ACL hydration so retrieval respects source permissions, OpenTelemetry traces). Design question for both stacks: where does the causal graph of the use case live before anyone tunes a prompt?

  • Measure

    Hallucination detection is not a checkbox. Recent theory maps automated hallucination detection to language identification in the Gold-Angluin sense: without explicit negative examples, reliable detection is impossible for most language collections. A separate impossibility result argues perfect hallucination control cannot satisfy truthfulness, information conservation, and knowledge revelation simultaneously. The practical read: stop selling detection as elimination; measure error budgets, human review load, and downstream harm.

    Counterweight worth watching: TabPFN and tabular foundation models. TabPFN-3 scales to far larger tables with a single forward pass and no hyperparameter tuning, topping TabArena on many benchmarks. For regulated and scientific workloads, the right measure is often still a calibrated tabular model, not another LLM judge stacked on top.

  • Decide

    Agent governance is a finance problem now. Per-token prices fell sharply since 2022, yet enterprise AI bills rose an estimated 320% as agentic workflows multiply steps per task. Reports cite Uber exhausting its 2026 AI coding budget by April, Microsoft pulling back broad Claude Code access, and engineers running $500–$2,000/month in token spend. Gartner placed generative AI in the trough of disillusionment and forecast 25% of 2026 AI budget slipping to 2027 as POCs stall. Productivity may be real; procurement modeled the wrong utility.

    LLM slop is the cultural mirror: Merriam-Webster made "slop" its 2025 Word of the Year for low-effort AI content at scale. The decide line for practitioners: ship artifacts with named evaluation, not volume. That is the bar Designed & Measured AI is written for.

Also on the radar

  • Agent token load: Jellyfish data cited in industry coverage puts per-developer consumption ~18× higher over nine months as coding agents go mainstream.
  • Identity-first agents: Directory-backed agent IDs (Okta + orchestration platforms) as the control plane for revoke and audit, not shared API keys.
  • Information ecosystem: Columbia IGP's March 2026 convening on AI slop and democratic information quality (research agenda, not hot take).
  • Real-time lakehouse: Lakehouse RT / Reyden as the substrate for in-the-moment agent reads without a separate serving tier.

Thank you, NC State IAA

Panelists and speakers at the inaugural NC State Institute for Advanced Analytics Alumni Summit
Inaugural IAA Alumni Summit, Talley Student Center, NC State.

Thank you to Dr. Rand, the NC State Institute for Advanced Analytics, April Wilson, Valerie Schwartz, and my fellow panelists for a great discussion at the inaugural Alumni Summit this spring. The event brought together alumni across nearly two decades of MSA cohorts to discuss how AI is reshaping analytics careers and organizations.

For anyone interested, NC State's recap is here:

We covered analytics careers, generative AI adoption, and why LLM productivity gains have been easier to observe in software development than in most forms of knowledge work.

One story stuck with me. I've increasingly found myself reviewing AI-generated strategy documents and wondering who actually owns the ideas. The writing is polished. The recommendations sound plausible. Yet the accountability, domain judgment, and decision-making rigor are often missing.

It reinforced something I've been thinking about for a while: as AI makes content abundant, value shifts toward evaluation, governance, and judgment.

Model costs will continue to fall. Capabilities will continue to improve. The constraint on AI impact is becoming less about model performance and more about how organizations make decisions.

The hardest part of AI is often not building the model. It's changing how decisions get made.

Commercial gen AI in life sciences

While at McKinsey QuantumBlack, I contributed to a May 2024 synthesis on generative AI in commercial life sciences. The underlying December 2023 survey of more than 100 pharma, biotech, and medtech leaders still reads cleanly for teams moving past pilots toward funded, custom build. What I took from those client conversations: adoption and budget were often ahead of decisions on evaluation, operating models, and who owns the output when a model reaches the field team.

McKinsey & Company: Early adoption of generative AI in commercial life sciences McKinsey & Company Early adoption of generative AI in commercial life sciences Life Sciences · May 2024 · Survey synthesis Read article (PDF) →