mirror of https://git.postgresql.org/git/postgresql.git synced 2026-02-22 14:27:00 +08:00

Go to file

John Naylor ef3c3cf6d0 Perform radix sort on SortTuples with pass-by-value Datums

Radix sort can be much faster than quicksort, but for our purposes it
is limited to sequences of unsigned bytes. To make tuples with other
types amenable to this technique, several features of tuple comparison
must be accounted for, i.e. the sort key must be "normalized":

1. Signedness -- It's possible to modify a signed integer such that
it can be compared as unsigned. For example, a signed char has range
-128 to 127. If we cast that to unsigned char and add 128, the range
of values becomes 0 to 255 while preserving order.

2. Direction -- SQL allows specification of ASC or DESC. The
descending case is easily handled by taking the complement of the
unsigned representation.

3. NULL values -- NULLS FIRST and NULLS LAST must work correctly.

This commmit only handles the case where datum1 is pass-by-value
Datum (possibly abbreviated) that compares like an ordinary
integer. (Abbreviations of values of type "numeric" are a convenient
counterexample.) First, tuples are partitioned by nullness in the
correct NULL ordering. Then the NOT NULL tuples are sorted with radix
sort on datum1. For tiebreaks on subsequent sortkeys (including the
first sort key if abbreviated), we divert to the usual qsort.

ORDER BY queries on pre-warmed buffers are up to 2x faster on high
cardinality inputs with radix sort than the sort specializations added
by commit 697492434, so get rid of them. It's sufficient to fall back
to qsort_tuple() for small arrays. Moderately low cardinality inputs
show more modest improvents. Our qsort is strongly optimized for very
low cardinality inputs, but radix sort is usually equal or very close
in those cases.

The changes to the regression tests are caused by under-specified sort
orders, e.g. "SELECT a, b from mytable order by a;". For unstable
sorts, such as our qsort and this in-place radix sort, there is no
guarantee of the order of "b" within each group of "a".

The implementation is taken from ska_byte_sort() (Boost licensed),
which is similar to American flag sort (an in-place radix sort) with
modifications to make it better suited for modern pipelined CPUs.

The technique of normalization described above can also be extended
to the case of multiple keys. That is left for future work (Thanks
to Peter Geoghegan for the suggestion to look into this area).

Reviewed-by: Chengpeng Yan <chengpeng_yan@outlook.com>
Reviewed-by: zengman <zengman@halodbtech.com>
Reviewed-by: ChangAo Chen <cca5507@qq.com>
Reviewed-by: Álvaro Herrera <alvherre@kurilemu.de>
Reviewed-by: Chao Li <li.evan.chao@gmail.com> (earlier version)
Discussion: https://postgr.es/m/CANWCAZYzx7a7E9AY16Jt_U3+GVKDADfgApZ-42SYNiig8dTnFA@mail.gmail.com

2026-02-14 13:50:06 +07:00

.github

Add CODE_OF_CONDUCT.md, CONTRIBUTING.md, and SECURITY.md.

2024-07-02 13:03:58 -05:00

config

Revert "Change copyObject() to use typeof_unqual"

2026-02-07 10:08:38 +01:00

contrib

Add support for INSERT ... ON CONFLICT DO SELECT.

2026-02-12 09:57:04 +00:00

doc

doc: Mention PASSING support for jsonpath variables

2026-02-13 12:12:11 +01:00

src

Perform radix sort on SortTuples with pass-by-value Datums

2026-02-14 13:50:06 +07:00

.cirrus.star

ci: Simplify ci-os-only handling

2025-08-14 12:09:34 -04:00

.cirrus.tasks.yml

ci: Configure g++ with 32-bit for 32-bit build

2026-01-09 08:58:50 +01:00

.cirrus.yml

ci: Per-repo configuration for manually trigger tasks

2025-08-14 11:54:03 -04:00

.dir-locals.el

Make Emacs perl-mode indent more like perltidy.

2019-01-13 11:32:31 -08:00

.editorconfig

Update .editorconfig and .gitattributes for postgresql.conf.sample.

2025-11-18 10:28:36 -06:00

.git-blame-ignore-revs

Add a couple of recent commits to .git-blame-ignore-revs.

2026-01-28 15:56:48 -06:00

.gitattributes

Update .editorconfig and .gitattributes for postgresql.conf.sample.

2025-11-18 10:28:36 -06:00

.gitignore

Update top-level .gitignore.

2022-12-04 15:23:00 -05:00

.mailmap

Add a Git .mailmap file

2024-11-05 13:56:02 +01:00

aclocal.m4

autoconf: Move export_dynamic determination to configure

2022-12-06 18:55:28 -08:00

configure

Revert "Change copyObject() to use typeof_unqual"

2026-02-07 10:08:38 +01:00

configure.ac

Revert "Change copyObject() to use typeof_unqual"

2026-02-07 10:08:38 +01:00

Update copyright for 2026

2026-01-01 13:24:10 -05:00

GNUmakefile.in

Allow selecting the git revision to be packaged by "make dist".

2024-05-03 11:08:50 -04:00

HISTORY

Canonicalize some URLs

2020-02-10 20:47:50 +01:00

Makefile

Remove AIX support

2024-02-28 15:17:23 +04:00

meson_options.txt

2026-01-01 13:24:10 -05:00

meson.build

meson: Add target for generating docs images

2026-02-13 11:50:14 +01:00

README.md

Revise the style of a paragraph in README.md.

2024-03-21 10:16:41 -05:00

README.md

PostgreSQL Database Management System

This directory contains the source code distribution of the PostgreSQL database management system.

PostgreSQL is an advanced object-relational database management system that supports an extended subset of the SQL standard, including transactions, foreign keys, subqueries, triggers, user-defined types and functions. This distribution also contains C language bindings.

General documentation about this version of PostgreSQL can be found at https://www.postgresql.org/docs/devel/. In particular, information about building PostgreSQL from the source code can be found at https://www.postgresql.org/docs/devel/installation.html.

The latest version of this software, and related software, may be obtained at https://www.postgresql.org/download/. For more information look at our web site located at https://www.postgresql.org/.

Languages

C 84.8%

PLpgSQL 6.1%

Perl 4.7%

Yacc 1.2%

Meson 0.7%

Other 2.4%