The Data Architecture Behind AI Agents
AWS, Google Cloud, Confluent, Snowflake, and Materialize are converging on the same requirement: agents need governed context that is fresh, stateful, and fast to retrieve.
ReadLatest
Open source analytics
DuckLabs joined AWS on September 1 after the acquisition process concluded on August 31. DuckDB and its related projects remain MIT-licensed under the independent DuckDB Foundation.
ReadFrom the blog
AWS, Google Cloud, Confluent, Snowflake, and Materialize are converging on the same requirement: agents need governed context that is fresh, stateful, and fast to retrieve.
ReadGlue 6.0, Snowpipe Streaming, Openflow, and cross-cloud BigQuery access show Iceberg moving beyond a file format—but interoperability still depends on catalogs, capabilities, and ownership.
ReadAirflow 3.0 adds assets, watchers, and event-driven scheduling, letting DAGs react to meaningful state changes—while preserving important limits around triggers, trust, and event semantics.
ReadPartition count is a throughput, parallelism, ordering, and recovery decision. AWS's new MSK guidance provides a useful sizing method—but measurement remains the final authority.
ReadSnowflake's open-source data-eng-bench tests agents inside real dbt projects and grades materialized data. Its strongest lesson applies beyond AI: benchmark the system's delivered result, not an isolated code snippet.
ReadPublic engineering reports from GitHub, Uber, Netflix, Airbnb, and AWS reveal practical ways to protect data integrity, recover state, drain backlogs, preserve observability, and test restoration.
ReadThe role began with authority over enterprise data problems. Cloud, analytics, regulation, and AI have expanded it into an operating mandate that connects trusted data to measurable outcomes.
ReadAmazon MSK Express can now materialize Kafka topics as read-only Iceberg tables, with AWS handling write coordination, delivery semantics, file sizing, and compaction.
ReadDatabricks is rolling row tracking and Checkpoint V2 onto eligible existing Unity Catalog tables—but only after a long compatibility observation window.
ReadGoogle announced GA task types for tables, views, sources, and quality tests, but its current pipeline guide still marks important parts of the workflow as Preview.
ReadAdaptive refresh can choose when to rebuild, while custom incrementalization lets engineers supply the MERGE or INSERT—two different answers to awkward change patterns.
ReadA Beta framework can exercise SQL, Auto CDC, streaming tables, expectations, and dependent transformations—provided every tested boundary is a catalog table name.
ReadGA Unity Catalog Python UDTFs give SQL callers a governed table-shaped interface to Python, with explicit runtime, dependency, and isolation limits.
ReadApache Ossie import and export, externally triggered hybrid jobs, and GA Cost Insights show dbt adapting to stacks it does not fully control.
ReadRuntime 19 is GA on Spark 4.2, but JDK, Python-package, UDF-serialization, and security changes make it a compatibility project—not a version-number flip.
Read