Skip to main content
Lambda pre-aggregations follow the Lambda architecture design to union real-time and batch data. Cube acts as a serving layer and uses pre-aggregations as a batch layer and source data or other pre-aggregations, usually streaming, as a speed layer. Due to this design, lambda pre-aggregations only work with data that is newer than the existing batched pre-aggregations.
Lambda pre-aggregations only work with Cube Store.

Use cases

Below we are looking at the most common examples of using lambda pre-aggregations.

Batch and source data

Batch data is coming from pre-aggregation and real-time data is coming from the data source.
Lambda pre-aggregation batch and source diagram
First, you need to create pre-aggregations that will contain your batch data. In the following example, we call it batch. Please note, it must have a time_dimension and partition_granularity specified. Cube will use these properties to union batch data with freshly-retrieved source data. You may also control the batch part of your data with the build_range_start and build_range_end properties of a pre-aggregation to determine a specific window for your batched data. Next, you need to create a lambda pre-aggregation. To do that, create pre-aggregation with type rollup_lambda, specify rollups you would like to use with rollups property, and finally set union_with_source_data: true to use source data as a real-time layer. Please make sure that the lambda pre-aggregation definition comes first when defining your pre-aggregations.

Batch and streaming data

In this scenario, batch data is comes from one pre-aggregation and real-time data comes from a streaming pre-aggregation.
Lambda pre-aggregation batch and streaming diagram
You can use lambda pre-aggregations to combine data from multiple pre-aggregations, where one pre-aggregation can have batch data and another streaming. Cube serves each date range with the first rollup in the list that has a fully built partition for it. Each next rollup is used only after the last partition served by the previous one. Partitions of the last rollup are used even if they are not completely built. Build ranges of the referenced rollups have to overlap, see below.

Overlapping build ranges

A partition that isn’t fully built is skipped for every rollup except the last one. When the build range of a rollup starts exactly where the previous one ends, the partition that has just become complete, for example the previous day right after midnight, is served by no rollup until it’s rebuilt. Such days are missing from query results without an error. To avoid that, start the build range of each rollup earlier than the build range of the previous one ends, with a margin longer than the previous rollup’s refresh_key interval plus its build time. Rows are not counted twice: a rollup skips partitions that the previous one already serves.
If a rollup uses a coarser partition_granularity than the next one, the next rollup has to start before the beginning of the previous rollup’s last partition, because that whole partition is skipped until it’s complete. For example, with month partitions on batch and day partitions on hot, start hot at date_trunc('month', CURRENT_DATE - INTERVAL '16 days').