your123 - stock.adobe.com

Streamhouse: A new way to serve AI with real-time context

A group of vendors is joining forces to define an architecture for fueling AI with real-time context to aid enterprises struggling to feed agents the streaming data they need.

Meet the streamhouse.

Created in response to the demand for AI agents to be served with the context they require to perform properly in production, this portmanteau of streaming data and data lakehouse builds on the data lakehouse foundation to address AI agents' need for continuous, real-time data.

Aiven, Confluent, Redpanda, StreamNative and Ververica recently established the Streamhouse Working Group and introduced an open definition of the new streamhouse architecture and a list of its key attributes.

Lakehouse architectures enable users to store and analyze large volumes of historical data. Streamhouse architectures similarly enable access to relevant data. But where streamhouses differ from lakehouses is in the data timeframe: rather than house historical data, streamhouses simplify and standardize serving real-time data spread across myriad systems to agents so they're informed by the fresh context that keeps agentic workflows current and reliable.

"Publishing a standardized architecture matters, even if only because it lays down a shared language and a common starting point for the industry," Donald Farmer, founder and principal of TreeHive Strategy, told TechTarget. "The streamhouse definition provides a vendor-neutral blueprint."

Kevin Petrie, an analyst at BARC U.S., similarly noted that defining a new architecture -- including naming supporting technologies such as change data capture, open table formats and data catalogs -- is valuable for the data and AI industries. However, he noted that the working definition of a streamhouse is nascent and specifics such as how the architecture supports other open standards are scarce.

"I fully support the concept and believe this working group calls out the right supporting technologies," he told TechTarget. "With that said, the Streamhouse website documentation is pretty sparse so far and would benefit from a more detailed reference architecture based on customer implementations."

A need for something new

Agents are designed to make decisions and execute business processes such as supply chain management and fraud detection as operational events take place.

To perform as intended in production, agents require not only high-quality data that can be trusted, but also fresh, governed business context delivered continuously and on a production scale. Without streaming data, they lack real-time context and can miss the triggers they need to carry out business processes as intended.

Publishing a standardized architecture matters, even if only because it lays down a shared language and a common starting point for the industry. The streamhouse definition provides a vendor-neutral blueprint.
Donald FarmerFounder and principal, TreeHive Strategy

However, enterprise data is often distributed across a morass of SaaS applications and stored in databases, data warehouses, data lakes and data lakehouses on multiple clouds. As a result, building pipelines that connect data spread across so many systems and delivering it to agents in real-time -- with the appropriate governance -- is too complex for some organizations, according to Farmer.

"The ad-hoc plumbing most enterprises had in place just wasn't good enough," he said.

A new way of serving agents the real-time data they require is therefore needed, Farmer continued.

"The streamhouse model specifically engineers streams for production environments, adding the necessary transformation and governance that ad-hoc messaging pipelines have not delivered," he said. "It also enables organizations to avoid the friction of consolidating all enterprise data into a single system before use."

Petrie pointed out that while streaming data is important for agents, as enterprises race to develop and deploy autonomous AI networks, not enough emphasis has been placed on how agents access real-time data.

"Amidst the rush to embrace agents, lakehouses, and documents, I think a lot of organizations overlooked the need to provide their agents with a standard, open method of accessing data on a real-time basis," he said.

Streaming architectures such as Apache Kafka, which is the technology Confluent is built on, offer one way of doing so, Petrie continued. The streamhouse is now another.

"Nearly all agents require real-time access to data, and streaming is an ideal way to provide real-time access and processing of tabular or semi-structured logs," Petrie said.

Aiven, Confluent, Redpanda, StreamNative and Ververica define a streamhouse as an architecture that enables enterprises to continuously capture, govern and serve the current state of their operations to agents and other applications so they can analyze and act upon the information in real time.

Key attributes of the streamhouse architecture include the following:

  • Continuous capture, processing and delivery of data so it is made available in real time as events occur rather than delivered via batch file processing.
  • Production-grade engineering so that business-critical applications, including agents, and key analyses can depend on continuous delivery.
  • Decentralization so agents can access data where it lives rather than wait for it to be packaged and consolidated in a single location.

Alex Gallego, co-founder and CEO of Redpanda, noted that the vendors came together to define the streamhouse after observing customers originate from different starting points to build agent workflows that ultimately looked similar, each featuring real-time data and production-grade engineering that connect decentralized data estates.

"The data can't be hours old, and the system can't depend on a batch job catching up before an agent makes a decision," Gallego told TechTarget. "We were all hearing versions of that from customers. When we compared notes, it was clear this … was the market converging on the same problem. That convergence is what made a shared, vendor-neutral definition worth publishing."

While the key attributes named by the Streamhouse Working Group are valuable, Farmer noted that more capabilities are also critical to developing an effective pipeline to serve agents with streaming data. In particular, governance, security and interoperability that enable streamhouses to exchange data with lakehouses using open table formats such as Apache Iceberg and Delta Lake are also vital.

"If an agent is taking autonomous action, the data feeding it must be strictly audited, lineage-tracked, and access-controlled," Farmer said. 

Petrie noted that the Streamhouse Working Group's correctly emphasizes certain governance capabilities. However, like Farmer, he advised the consortium to add more specifics regarding how streamhouses can keep streaming data secure and compliant with evolving regulations.

"The working group rightly names governance as a must-have characteristic, but I think users need more detail about how best to meet specific requirements for identity and access management, personally identifiable information obfuscation, and protection of intellectual property," he said. "They also need more information about how catalogs and meta stores can organize the supporting metadata."

The streamhouse needs additions

After publishing the initial definition of a streamhouse, the Streamhouse Working Group plans to maintain the definition and alter it as technology evolves, according to Gallego.

In addition, the group is working to improve the standards and practices that allow organizations to build streamhouses using both open-source and proprietary tooling and opening participation in the Streamhouse Working Group to additional members.

"Streamhouse is open by design," Gallego said. "No one company owns the definition, and we welcome anyone building in this space to contribute and be part of it. That's better for customers -- they get choice, competition, and innovation instead of another category controlled by one vendor."

Gallego acknowledged that the streamhouse architecture addresses only part of the problem that prevents organizations from putting agents in production. While the architecture provides the data infrastructure, a governance and security layer is still needed to control what agents access.

An additional missing piece that Farmer advised the working group to broach is how streamhouses interface with the memory systems of agents, including vector databases and retrieval-augmented generation systems, so data is prepared in real time.

"Transporting the data is only half the battle," he said. "Formatting and indexing the streaming data so an LLM can query it -- in the milliseconds needed for real-time -- is a critical hurdle they do not address."

Eric Avidon is a senior news writer for Informa TechTarget and a journalist with more than three decades of experience. He covers analytics and data management.

Dig Deeper on Data Management