Designed & Measured AI
Reading
A working shelf for enterprise AI: who to follow, where to listen, which papers to keep nearby, and a feed sorted through Design · Measure · Decide. Updated by a local fetch script; LinkedIn-only voices are listed manually.
Who I'm following
Practitioners whose public work shapes how enterprise AI gets designed, measured, and decided. Priority sources from the benchmark cohort; LinkedIn profiles are manual follows (no RSS).
-
Sohrab Rahimi
DesignPartner, QuantumBlack / AI by McKinsey
Primary reference for the McKinsey-to-enterprise arc. Long-form LinkedIn posts that synthesize agentic AI papers into structural failure modes and numbered frameworks. The paper-to-implication structure is the template for Designed & Measured carousels.
LinkedIn · Manual follow (no RSS)
-
Andriy Burkov
GeneralAuthor, The Hundred-Page Machine Learning Book
Educational carousels and accessible technical depth at high cadence. Useful for format calibration: beginner-friendly without dumbing down, consistent posting rhythm.
-
Eugene Yan
MeasureApplied ML in production
Production ML honesty, recsys, and applied LLM patterns. Reading-list curation and "what actually works in prod" framing.
Writing · RSS in feed
-
Chip Huyen
DesignAI engineering, ML systems
Newsletter-to-book-to-stage pipeline. AI Engineering as the credibility anchor for practitioner depth that ages well.
Blog · RSS in feed
-
Hamel Husain
MeasureLLM evals, Parlance Labs
Eval-first practitioner posts. The reference for how to measure agentic workflows without theater metrics.
Newsletter · RSS in feed
-
Lilian Weng
DesignLil'Log · OpenAI
Deep technical writing on agents, alignment, and ML foundations. Long-form reference when a topic needs first principles.
Lil'Log · RSS in feed
-
Cassie Kozyrkov
DecideDecision intelligence
Named frameworks for executives without losing practitioner respect. Model for the Four-Part ROI Test and executive "so what" framing.
-
Swyx (Shawn Wang)
DesignLatent Space · AI Engineer Summit
Podcast and community as career accelerator. Category ownership through consistent technical interviews.
Latent Space · RSS in feed
Podcasts & newsletters
Shows and publications that stay current on enterprise AI, evals, and infrastructure. Items from RSS-capable sources appear in the live feed below.
-
Everyday AI
DecideDaily podcast and newsletter from Jordan Wilson on practical AI adoption for business leaders. Strong on trends, tooling, and workplace rollout (not deep evals).
Episodes · Auto-fetch RSS
-
KDnuggets
GeneralGregory Piatetsky-Shapiro's long-running hub for data science, ML, and AI news. Broad scan for tools, tutorials, and industry shifts.
Site · Auto-fetch RSS
-
Latent Space
DesignTop-tier AI engineering podcast and newsletter. Technical deep dives with the people behind major projects.
Site · Auto-fetch RSS
-
TWIML AI Podcast
GeneralSam Charrington interviews researchers and practitioners across ML and AI. Reliable breadth for signal on what's moving.
Site · Auto-fetch RSS
-
The Data Exchange
GeneralBen Lorica on data infrastructure, applied ML, and AI engineering. Strong for platform and production context.
Site · Auto-fetch RSS
-
Prior Labs
MeasureTabPFN and tabular foundation models. Counterweight to LLM-everything hype in regulated and scientific workloads.
Site · Auto-fetch RSS
-
Google AI Blog
GeneralResearch releases and applied AI from Google. Useful for benchmark and capability shifts worth tracking.
Site · Auto-fetch RSS
-
Designed & Measured AI
DecideOwn publication. Carousel depth expanded into essays on design, measurement, and decision quality for enterprise AI.
Papers worth tracking
Seed shelf from the LinkedIn carousel plan and reference reading. Anchors for hallucination, scientific AI, and tabular foundations.
-
Theoretical argument that calibration and zero hallucination are incompatible for language models. Core reference for Post 2 (hallucination detection limits).
-
Maps hallucination detection to language identification in the Gold-Angluin sense. Practical read: measure error budgets, not elimination claims.
-
Sampling-based consistency check without external KB. Useful baseline; production limits matter for enterprise guardrails.
-
Reference-free RAG evaluation metrics. Know where the framework stops (faithfulness vs. downstream harm).
-
Tabular foundation model with strong small-data performance. Counterweight to stacking LLM judges on regulated workloads.
-
Neural network potential for organic molecules. Formulation AI reference from Post 3 (scientific AI depth).
-
Structure prediction for proteins, nucleic acids, and ligands. Track capabilities and limits for life sciences AI narratives.
-
Molecular embeddings from masked language modeling. Baseline for formulation and cheminformatics stacks.
-
Benchmark contamination and data-poisoning practicality. Relevant to eval trust and training-data hygiene.
-
How reliable are LLM judges? Foundational for clinical and enterprise eval design.
-
Agent memory and planning architecture. Useful contrast to production agent stacks (identity, ACLs, observability).
-
Structured risk framing for deployed assistants. Governance and procurement language for enterprise rollouts.
-
Metric choice can create apparent emergent abilities. Skepticism anchor for benchmark hype.
-
Design
GPT-4 Technical Report
Baseline capability reference. Track what changed (and what did not) in successor models and enterprise expectations.
-
Legal and procedural framing for eval and red teaming. Relevant as agent governance matures.
Recent from the feed
Auto-fetched from RSS sources in scripts/reading_sources.yaml. Each item is tagged Design, Measure, Decide, or General via keyword heuristics. Run the fetch script weekly to refresh; sync copies JSON to website/data/reading-feed.json for deploy.
Loading feed…