Skip to content

Common subexpression elimination in columnar agg - #9534

Merged
akuzm merged 26 commits into
timescale:mainfrom
akuzm:cse
Jun 25, 2026
Merged

Common subexpression elimination in columnar agg#9534
akuzm merged 26 commits into
timescale:mainfrom
akuzm:cse

Conversation

@akuzm

@akuzm akuzm commented Apr 7, 2026

Copy link
Copy Markdown
Member
  1. Use interning to find the identical expression subtrees and replace
    them with a canonical pointer.
  2. Count how many times each expression is present in the hierarchy.
  3. For expression that are present many times, compute them only once
    per batch.

akuzm added 2 commits April 7, 2026 19:24
1) Use interning to find the identical expression subtrees and replace
   them with a canonical pointer.
1) Count how many times each expression is present in the hierarchy.
1) For expression that are present many times, compute them only once
   per batch.
@codecov

codecov Bot commented Apr 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@akuzm
akuzm marked this pull request as ready for review June 12, 2026 12:24
@github-actions
github-actions Bot requested review from antekresic and dbeck June 12, 2026 12:24
@github-actions

Copy link
Copy Markdown

@antekresic, @dbeck: please review this pull request.

Powered by pull-review

Comment thread tsl/test/expected/vector_agg_expr_debug.out
Comment thread tsl/test/sql/vector_agg_expr.sql

@antekresic antekresic left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved with a suggestion, otherwise lgtm

Comment thread tsl/src/nodes/vector_agg/exec.c
@akuzm
akuzm enabled auto-merge (squash) June 25, 2026 10:07
@akuzm
akuzm merged commit 81baae6 into timescale:main Jun 25, 2026
63 of 66 checks passed
@akuzm
akuzm deleted the cse branch June 25, 2026 10:17
kpan2034 pushed a commit to kpan2034/timescaledb that referenced this pull request Jul 1, 2026
1) Use interning to find the identical expression subtrees and replace
   them with a canonical pointer.
1) Count how many times each expression is present in the hierarchy.
1) For expression that are present many times, compute them only once
   per batch.
@surister surister mentioned this pull request Jul 7, 2026
@surister surister mentioned this pull request Jul 21, 2026
surister pushed a commit that referenced this pull request Jul 28, 2026
## 2.29.0 (2026-07-28)

This release contains performance improvements and bug fixes since the
2.28.3 release. We recommend that you upgrade at the next available
opportunity.

**Release Highlights**
* **Chunk exclusion for DML operations** drastically improves the
performance of `UPDATE` and `DELETE` statements on hypertables. By
acquiring exclusive locks only on the specific chunks being modified
rather than the entire hypertable, this enhancement eliminates massive
lock contention and keeps high-concurrency workloads running smoothly
without unnecessary slowdowns.
* Intelligent **row-by-row decompression** enables the query planner to
decompress data row-by-row rather than in large batches when an
operation prioritizes a fast initial response (such as queries with
`LIMIT` clauses). This dramatically reduces memory overhead and query
latency, ensuring lightning-fast performance when you only need to
retrieve a small subset of records from your compressed hypertables.

**Important: PostgreSQL 15 Support Removed**
TimescaleDB 2.29.0 removes support for PostgreSQL 15. This release
supports PostgreSQL 16, 17, and 18. If you are still running PostgreSQL
15, upgrade PostgreSQL before upgrading to TimescaleDB 2.29.0.

**Backward-Incompatible Changes**
* [#10041](#10041) Remove
support for PostgreSQL 15

**Features**
* [#9315](#9315) Speed up
`DML` operations on hypertables by using the optimized TimescaleDB
hypertable expansion code instead of the generic PostgreSQL inheritance
hierarchy expansion
* [#9534](#9534) Speed up
expression evaluation in the columnar pipeline by caching common
subexpressions
* [#9684](#9684) Add
`_timescaledb_functions.decompress_batch()` SQL function
* [#9732](#9732) Speed up
some queries with small `LIMIT` by switching to row-by-row query
execution pipeline
* [#9917](#9917) Decompress
less data in `DML` on compressed hypertables by accounting for prepared
statement parameters
* [#9957](#9957) Add
`compact_chunk()` function
* [#10048](#10048) Support
concurrent refresh policies on hierarchical continuous aggregates
* [#10081](#10081) Add
`samplerate` argument to
`_timescaledb_functions.estimate_uncompressed_size()`
* [#10100](#10100) Skip
classifying compressed relations to speed up planning
* [#10118](#10118) Don't
track compressed relations as separate chunk
* [#10119](#10119) Reduce
memory usage of `INSERT` queries using direct compress and spanning
multiple chunks
* [#10163](#10163) Add a
compaction policy for unordered chunks
* [#10204](#10204) Don't
create separate hypertable catalog entry for hypertables with
compression
* [#10217](#10217) Initial
placeholder version of granular refresh API
* [#10225](#10225) Add
`config_merge` parameter to `alter_job()` for merging `jsonb` into the
existing job configuration
* [#10226](#10226) Add
`recompress_unordered` columnstore policy option
* [#10231](#10231) Use
`regclass` for storing relation reference in chunk table
* [#10237](#10237) Add
helper functions for decoding hypertable status
* [#10240](#10240) Add the
`tsdb.direct_compress` storage parameter that allows enabling direct
compress for a given hypertable independent of global settings
* [#10266](#10266) Add
`max_batches` to `compact_chunk()`
* [#10299](#10299) Add
`continuous_aggs_tenant_tracking` and `hypertable_cagg_settings`
catalogs

**Bugfixes**
* [#10013](#10013) Make
ownership error messages on continuous aggregates consistent
* [#10052](#10052) Result
of `MIN` / `MAX` aggregate functions in columnar aggregation pipeline
possibly inconsistent with plain PostgreSQL result
* [#10071](#10071) Prune
the real-time branch of hierarchical continuous aggregates at any
nesting depth
* [#10143](#10143) Fix
division by zero when planning `time_bucket` with zero width
* [#10199](#10199) Fix
`initial_start` handling in `build_job_info`
* [#10213](#10213) Cache
sort pathkeys per hypertable
* [#10221](#10221) Fix
incremental refresh skipping the last bucket
* [#10278](#10278) Drop
`job_errors` view in `bgw_job_stat_history` migration
* [#10280](#10280)
`RETURNING` clause returned no rows for `INSERT` using direct compress
* [#10281](#10281) Disable
direct compress when the destination table has an exclusion constraint
so the constraint is still enforced
* [#10282](#10282) Only
count directly compressed rows toward the command tag when the `INSERT`
sets it
* [#10286](#10286)
Propagate `VACUUM` on a chunk to the compressed relation when running on
the chunk directly
* [#10302](#10302) Fix
useless-join removal and self-join elimination for hypertables
* [#10313](#10313) Allow
running `ALTER EXTENSION timescaledb UPDATE` inside a transaction block
* [#10315](#10315) Fix
overlap detection with running max
* [#10324](#10324) Fix
stale index entries after `rebuild_sparse_index()` on compressed chunks

**GUCs**
* `timescaledb.enable_hypertable_expansion_for_dml`: allow using the
optimized TimescaleDB hypertable expansion code for `UPDATE` and
`DELETE` instead of the generic PostgreSQL inheritance hierarchy
expansion. On by default.

**Thanks**
* @FrancescEthon and @ManuelEthon for reporting the issue
* @h0rn3t for reporting a problem with `VACUUM` not propagating to the
compressed relation
* @MaximeEthon for reporting an issue with prepared statement parameters
in DML decompression
* @proddata for reporting a problem when upgrading from 2.15.3 to 2.28.2
* @tureba for reporting and fixing stale sparse-index entries after
rebuild
* @viniciusrsouza for reporting an issue with hierarchical continuous
aggregates
svenklemm pushed a commit that referenced this pull request Jul 28, 2026
This release contains performance improvements and bug fixes since the
2.28.3 release. We recommend that you upgrade at the next available
opportunity.

**Release Highlights**
* **Chunk exclusion for DML operations** drastically improves the
performance of `UPDATE` and `DELETE` statements on hypertables. By
acquiring exclusive locks only on the specific chunks being modified
rather than the entire hypertable, this enhancement eliminates massive
lock contention and keeps high-concurrency workloads running smoothly
without unnecessary slowdowns.
* Intelligent **row-by-row decompression** enables the query planner to
decompress data row-by-row rather than in large batches when an
operation prioritizes a fast initial response (such as queries with
`LIMIT` clauses). This dramatically reduces memory overhead and query
latency, ensuring lightning-fast performance when you only need to
retrieve a small subset of records from your compressed hypertables.

**Important: PostgreSQL 15 Support Removed**
TimescaleDB 2.29.0 removes support for PostgreSQL 15. This release
supports PostgreSQL 16, 17, and 18. If you are still running PostgreSQL
15, upgrade PostgreSQL before upgrading to TimescaleDB 2.29.0.

**Backward-Incompatible Changes**
* [#10041](#10041) Remove
support for PostgreSQL 15

**Features**
* [#9315](#9315) Speed up
`DML` operations on hypertables by using the optimized TimescaleDB
hypertable expansion code instead of the generic PostgreSQL inheritance
hierarchy expansion
* [#9534](#9534) Speed up
expression evaluation in the columnar pipeline by caching common
subexpressions
* [#9684](#9684) Add
`_timescaledb_functions.decompress_batch()` SQL function
* [#9732](#9732) Speed up
some queries with small `LIMIT` by switching to row-by-row query
execution pipeline
* [#9917](#9917) Decompress
less data in `DML` on compressed hypertables by accounting for prepared
statement parameters
* [#9957](#9957) Add
`compact_chunk()` function
* [#10048](#10048) Support
concurrent refresh policies on hierarchical continuous aggregates
* [#10081](#10081) Add
`samplerate` argument to
`_timescaledb_functions.estimate_uncompressed_size()`
* [#10100](#10100) Skip
classifying compressed relations to speed up planning
* [#10118](#10118) Don't
track compressed relations as separate chunk
* [#10119](#10119) Reduce
memory usage of `INSERT` queries using direct compress and spanning
multiple chunks
* [#10163](#10163) Add a
compaction policy for unordered chunks
* [#10204](#10204) Don't
create separate hypertable catalog entry for hypertables with
compression
* [#10217](#10217) Initial
placeholder version of granular refresh API
* [#10225](#10225) Add
`config_merge` parameter to `alter_job()` for merging `jsonb` into the
existing job configuration
* [#10226](#10226) Add
`recompress_unordered` columnstore policy option
* [#10231](#10231) Use
`regclass` for storing relation reference in chunk table
* [#10237](#10237) Add
helper functions for decoding hypertable status
* [#10240](#10240) Add the
`tsdb.direct_compress` storage parameter that allows enabling direct
compress for a given hypertable independent of global settings
* [#10266](#10266) Add
`max_batches` to `compact_chunk()`
* [#10299](#10299) Add
`continuous_aggs_tenant_tracking` and `hypertable_cagg_settings`
catalogs

**Bugfixes**
* [#10013](#10013) Make
ownership error messages on continuous aggregates consistent
* [#10052](#10052) Result
of `MIN` / `MAX` aggregate functions in columnar aggregation pipeline
possibly inconsistent with plain PostgreSQL result
* [#10071](#10071) Prune
the real-time branch of hierarchical continuous aggregates at any
nesting depth
* [#10143](#10143) Fix
division by zero when planning `time_bucket` with zero width
* [#10199](#10199) Fix
`initial_start` handling in `build_job_info`
* [#10213](#10213) Cache
sort pathkeys per hypertable
* [#10221](#10221) Fix
incremental refresh skipping the last bucket
* [#10278](#10278) Drop
`job_errors` view in `bgw_job_stat_history` migration
* [#10280](#10280)
`RETURNING` clause returned no rows for `INSERT` using direct compress
* [#10281](#10281) Disable
direct compress when the destination table has an exclusion constraint
so the constraint is still enforced
* [#10282](#10282) Only
count directly compressed rows toward the command tag when the `INSERT`
sets it
* [#10286](#10286)
Propagate `VACUUM` on a chunk to the compressed relation when running on
the chunk directly
* [#10302](#10302) Fix
useless-join removal and self-join elimination for hypertables
* [#10313](#10313) Allow
running `ALTER EXTENSION timescaledb UPDATE` inside a transaction block
* [#10315](#10315) Fix
overlap detection with running max
* [#10324](#10324) Fix
stale index entries after `rebuild_sparse_index()` on compressed chunks

**GUCs**
* `timescaledb.enable_hypertable_expansion_for_dml`: allow using the
optimized TimescaleDB hypertable expansion code for `UPDATE` and
`DELETE` instead of the generic PostgreSQL inheritance hierarchy
expansion. On by default.

**Thanks**
* @FrancescEthon and @ManuelEthon for reporting an issue with incremental refresh skipping the last bucket
* @h0rn3t for reporting a problem with `VACUUM` not propagating to the
* @igor2x for reporting an issue with locking during DML statements on hypertables
compressed relation
* @MaximeEthon for reporting an issue with prepared statement parameters
in DML decompression
* @proddata for reporting a problem when upgrading from 2.15.3 to 2.28.2
* @tureba for reporting and fixing stale sparse-index entries after
rebuild
* @viniciusrsouza for reporting an issue with hierarchical continuous
aggregates
surister pushed a commit to surister/timescaledb that referenced this pull request Jul 28, 2026
This release contains performance improvements and bug fixes since the
2.28.3 release. We recommend that you upgrade at the next available
opportunity.

**Release Highlights**
* **Chunk exclusion for DML operations** drastically improves the
performance of `UPDATE` and `DELETE` statements on hypertables. By
acquiring exclusive locks only on the specific chunks being modified
rather than the entire hypertable, this enhancement eliminates massive
lock contention and keeps high-concurrency workloads running smoothly
without unnecessary slowdowns.
* Intelligent **row-by-row decompression** enables the query planner to
decompress data row-by-row rather than in large batches when an
operation prioritizes a fast initial response (such as queries with
`LIMIT` clauses). This dramatically reduces memory overhead and query
latency, ensuring lightning-fast performance when you only need to
retrieve a small subset of records from your compressed hypertables.

**Important: PostgreSQL 15 Support Removed**
TimescaleDB 2.29.0 removes support for PostgreSQL 15. This release
supports PostgreSQL 16, 17, and 18. If you are still running PostgreSQL
15, upgrade PostgreSQL before upgrading to TimescaleDB 2.29.0.

**Backward-Incompatible Changes**
* [timescale#10041](timescale#10041) Remove
support for PostgreSQL 15

**Features**
* [timescale#9315](timescale#9315) Speed up
`DML` operations on hypertables by using the optimized TimescaleDB
hypertable expansion code instead of the generic PostgreSQL inheritance
hierarchy expansion
* [timescale#9534](timescale#9534) Speed up
expression evaluation in the columnar pipeline by caching common
subexpressions
* [timescale#9684](timescale#9684) Add
`_timescaledb_functions.decompress_batch()` SQL function
* [timescale#9732](timescale#9732) Speed up
some queries with small `LIMIT` by switching to row-by-row query
execution pipeline
* [timescale#9917](timescale#9917) Decompress
less data in `DML` on compressed hypertables by accounting for prepared
statement parameters
* [timescale#9957](timescale#9957) Add
`compact_chunk()` function
* [timescale#10048](timescale#10048) Support
concurrent refresh policies on hierarchical continuous aggregates
* [timescale#10081](timescale#10081) Add
`samplerate` argument to
`_timescaledb_functions.estimate_uncompressed_size()`
* [timescale#10100](timescale#10100) Skip
classifying compressed relations to speed up planning
* [timescale#10118](timescale#10118) Don't
track compressed relations as separate chunk
* [timescale#10119](timescale#10119) Reduce
memory usage of `INSERT` queries using direct compress and spanning
multiple chunks
* [timescale#10163](timescale#10163) Add a
compaction policy for unordered chunks
* [timescale#10204](timescale#10204) Don't
create separate hypertable catalog entry for hypertables with
compression
* [timescale#10217](timescale#10217) Initial
placeholder version of granular refresh API
* [timescale#10225](timescale#10225) Add
`config_merge` parameter to `alter_job()` for merging `jsonb` into the
existing job configuration
* [timescale#10226](timescale#10226) Add
`recompress_unordered` columnstore policy option
* [timescale#10231](timescale#10231) Use
`regclass` for storing relation reference in chunk table
* [timescale#10237](timescale#10237) Add
helper functions for decoding hypertable status
* [timescale#10240](timescale#10240) Add the
`tsdb.direct_compress` storage parameter that allows enabling direct
compress for a given hypertable independent of global settings
* [timescale#10266](timescale#10266) Add
`max_batches` to `compact_chunk()`
* [timescale#10299](timescale#10299) Add
`continuous_aggs_tenant_tracking` and `hypertable_cagg_settings`
catalogs

**Bugfixes**
* [timescale#10013](timescale#10013) Make
ownership error messages on continuous aggregates consistent
* [timescale#10052](timescale#10052) Result
of `MIN` / `MAX` aggregate functions in columnar aggregation pipeline
possibly inconsistent with plain PostgreSQL result
* [timescale#10071](timescale#10071) Prune
the real-time branch of hierarchical continuous aggregates at any
nesting depth
* [timescale#10143](timescale#10143) Fix
division by zero when planning `time_bucket` with zero width
* [timescale#10199](timescale#10199) Fix
`initial_start` handling in `build_job_info`
* [timescale#10213](timescale#10213) Cache
sort pathkeys per hypertable
* [timescale#10221](timescale#10221) Fix
incremental refresh skipping the last bucket
* [timescale#10278](timescale#10278) Drop
`job_errors` view in `bgw_job_stat_history` migration
* [timescale#10280](timescale#10280)
`RETURNING` clause returned no rows for `INSERT` using direct compress
* [timescale#10281](timescale#10281) Disable
direct compress when the destination table has an exclusion constraint
so the constraint is still enforced
* [timescale#10282](timescale#10282) Only
count directly compressed rows toward the command tag when the `INSERT`
sets it
* [timescale#10286](timescale#10286)
Propagate `VACUUM` on a chunk to the compressed relation when running on
the chunk directly
* [timescale#10302](timescale#10302) Fix
useless-join removal and self-join elimination for hypertables
* [timescale#10313](timescale#10313) Allow
running `ALTER EXTENSION timescaledb UPDATE` inside a transaction block
* [timescale#10315](timescale#10315) Fix
overlap detection with running max
* [timescale#10324](timescale#10324) Fix
stale index entries after `rebuild_sparse_index()` on compressed chunks

**GUCs**
* `timescaledb.enable_hypertable_expansion_for_dml`: allow using the
optimized TimescaleDB hypertable expansion code for `UPDATE` and
`DELETE` instead of the generic PostgreSQL inheritance hierarchy
expansion. On by default.

**Thanks**
* @FrancescEthon and @ManuelEthon for reporting an issue with incremental refresh skipping the last bucket
* @h0rn3t for reporting a problem with `VACUUM` not propagating to the
* @igor2x for reporting an issue with locking during DML statements on hypertables
compressed relation
* @MaximeEthon for reporting an issue with prepared statement parameters
in DML decompression
* @proddata for reporting a problem when upgrading from 2.15.3 to 2.28.2
* @tureba for reporting and fixing stale sparse-index entries after
rebuild
* @viniciusrsouza for reporting an issue with hierarchical continuous
aggregates
@timescale-automation timescale-automation added the released-2.29.0 Released in 2.29.0 label Jul 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

released-2.29.0 Released in 2.29.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants