The Daily Diff

The Daily Diff

by Premchand Chidipoti
Season 1

Netflix's Real-Time Distributed Graph, Part 1 — Ingesting and Processing Data Streams

AI
The origin story of Netflix's Real-Time Distributed Graph: why microservices left them with data silos across ads, live events, and games, and how member actions become graph nodes and edges. Jordan and Riley cover the API Gateway to Kafka to Flink pipeline (~1M msgs/sec per topic, 5M+ records/sec published downstream), and the two scaling calls Netflix made — one Flink job per Kafka topic, and a separate topic per node/edge type — trading operational overhead for independent scaling. Source: How and Why Netflix Built a Real-Time Distributed Graph, Part 1: Ingesting and Processing Data Streams at Internet Scale — Netflix Tech Blog, Oct 17 2025 — https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-1-ingesting-and-processing-data-80113e124acc This is commentary/summary in the hosts' own words, not a reproduction of the article.

Pilot Ep 1 — Cloudflare's Town Lake: A Data Platform, and the AI Agent Built on It

AI
In this episode, Jordan and Riley break down a recent Cloudflare Engineering post on Town Lake, their internal data lakehouse, and Skipper, the AI agent built on top of it. They cover why Cloudflare needed it (a billion+ events/sec scattered across Postgres, ClickHouse, Kafka, and more), how Trino and Iceberg tie it together into one queryable system, the "default-closed" access model for sensitive data, and how Skipper turns natural-language questions into safe, grounded SQL — including a neat trick that cut its tool-calling round-trips from five down to one. Source: How we built Cloudflare's data platform and an AI agent on top of it — The Cloudflare Blog, May 28, 2026. This episode is commentary and discussion, not a reproduction of the original post — read the full article at the link above.

Google Research: Cardiometabolic Risk from Smartphone Photos (PhotoScan)

A health-AI piece with a clever engineering core. Jordan and Riley cover PhotoScan, a deep-learning framework that estimates 3D body-composition metrics — body fat %, android-to-gynoid (apple vs. pear) fat ratio, and visceral-to-subcutaneous fat ratio — from ordinary 2D smartphone photos, to flag insulin resistance (which precedes type 2 diabetes by years and is poorly captured by BMI). The standout trick solves a data problem: pre-train a ResNet-50 (ImageNet-init) on UK Biobank (N=35,323) using 2D projections rendered from 3D MRI with DXA as ground truth, fuse image features with sex/height/weight/BMI, and output probability density functions (uncertainty, not point guesses); then fine-tune on real smartphone photos (PhotoBIA, N=677, with landmark detection picking best frames from 360-degree video) and validate on an independent cohort (N=132), hitting near-DXA agreement and beating smartwatch impedance sensors — while unlocking ratios impedance can't measure. Transferable lessons: bootstrap a cheap deployment modality from a data-rich hard-to-collect one via projection; fuse image + tabular; predict distributions when stakes are clinical. Caveat: investigational, not a cleared medical device. Source: Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery — Google Research Blog, Aug 17 2026 (paper: arXiv:2603.27017) — https://research.google/blog/seeing-beyond-bmi-estimating-cardiometabolic-risk-with-smartphone-imagery/ This is commentary/summary in the hosts' own words, not a reproduction of the article.
3 of 3