doris

Author	SHA1	Message	Date
wangbo	20e2d2e2f8	[Fix](executor)Fix workload thread start failed when follower convert to master	2024-05-12 09:30:14 +08:00
minghong	9915862bf7	[opt](nereids)estimate rowcount for is-null filter when column stats are not available (#34519 ) * estimate rowcount for is-null filter when column stats are not available	2024-05-11 15:04:35 +08:00
Mryange	719e50f353	[fix](json function) fix failed when json_exists_path use not null input (#34289 )	2024-05-11 15:04:35 +08:00
Pxl	1ff4dc8f85	[Bug](runtime-filter) fix coredump won change_null_to_true when argument column is not null… (#34602 ) fix coredump won change_null_to_true when argument column is not nullable	2024-05-11 15:04:35 +08:00
Kaijie Chen	2392477f76	[test](shuffle) test insert row count when rows filtered by ExchangeNode (#34657 )	2024-05-11 11:47:49 +08:00
HappenLee	8c237e82a3	[Bug](exec) fix intersections/differences bug (#34675 )	2024-05-11 11:45:31 +08:00
zhiqiang	58c19e33b3	[fix](round) Fix incorrect decimal scale inference in round functions (#34471 ) * FIX NEEDED * FORMAT * FORMAT * FIX TEST	2024-05-11 11:42:12 +08:00
HHoflittlefish777	7ba66c5890	[branch-2.1](routine-load) do not schedule task when there is no data (#34654 )	2024-05-11 11:01:18 +08:00
minghong	dd1b54cf62	[pick](nereids)Runtime filter pushdown refactor for branch-2.1 (#34682 ) * [refactor](Nereids)refactor runtime filter generator (#34275) 1. unify the process of generating rf for hash join and for nested loop join 2. fix some bugs in generating rf 3. remove some duplicated check (cherry picked from commit 07267faac0d9c6ef3bb1fd4ee101b4c761c8a2f2) * [refactor](nereids) do not deny a runtime filter by removing an entry in aliasMap (#34559) in current version, there are 2 approaches to verify whether a join condition can be used to generate a runtime filter, they are 1. remove the output slot from aliasMap 2. pushDownVisitor.visit(...) return false the 1st approach has some drawbacks, we prefer to the 2ed approach. In this pr, all the cases are handled by the 2ed approach, and remove the related code for the 1st approach. (cherry picked from commit a29082bf31e66efa2df193b38347e610f2bf7464) * rebase	2024-05-11 09:44:24 +08:00
yongjinhou	e38801968d	[Fix](functions) Fix bug in makedate and str_to_date functions	2024-05-10 22:14:25 +08:00
minghong	d5d6c7f8a4	[opt](nereids) optimize str-like-col range filter estimation (#34542 ) we have an order reserved mappping from string to double. for string column A, we have double values for A.min and A.max. when estimating A<"abc", A.min/max could be used to judge whether 'abc' is between A.min and A.max, but it cannot be used to do range estimation. suppose "abc" is mapped to double x. if we compute selectivity by formula "sel = (x-A.min)/(A.max-A.min)", we are likely to obtain extreme values.	2024-05-10 22:14:00 +08:00
morrySnow	845732b440	[WIP](test) remove enable_nereids_planner in regression cases (part 3) (#34558 ) previous PR: part 1: #34417 part 2: #34490	2024-05-10 22:11:01 +08:00
lihangyu	853dbdcb00	[Feature](PreparedStatement) implement general server side prepared (#33807 )	2024-05-10 22:10:11 +08:00
feiniaofeiafei	6c11dd2231	[Fix](planner) fix ScalarType.getAssignmentCompatibleType() when deal boolean and decimal (#34435 ) The legacy planner encounters issues when handling filters such as: c1(boolean type)=0.0(decimalv3). The literal 0.0 is interpreted as decimalv3(1,1), and the boolean type c1 is coerced to decimalv3(1,1). decimalv3(1,1) can only retain values in the range [0,1), while the boolean true is represented as 1, exceeding the upper bound, thus causing an overflow problem. This pull request addresses this issue by considering the boolean type as decimalv3(1,0), making both c1 and 0.0 being cast to decimal(2,1). Co-authored-by: feiniaofeiafei <moailing@selectdb.com>	2024-05-10 22:07:16 +08:00
feiniaofeiafei	7c56c17ecc	[Fix](nereids) fix NormalizeRepeat, change the outputExpression rewrite logic (#34196 ) In NormalizeRepeat, three parts of the outputExpression of LogicalRepeat need to be pushed down and outputted by bottom project: flattenGroupingSetExpr, argumentsOfGroupingScalarFunction, argumentsOfAggregateFunction. In the original code, use these three parts to rewrite the outputExpressions of LogicalRepeat to slots.This can cause problems in some cases, for example: ```sql SELECT ROUND( SUM(pk + 1) - 3) col_alias1, pk + 1 AS col_alias3 FROM table_20_undef_partitions2_keys3_properties4_distributed_by53 GROUP BY GROUPING SETS ((pk), ()) ; ``` The three parts expression needed to be pushed down are: pk, pk+1. The original code use pk+1 to rewrite the pk + 1 AS col_alias3 to slot. But the pk+1 is not in the list of grouping outputs, and then report error. This pr change the rewrite process, divide the expression needed to be pushed down into 2 parts: one is (flattenGroupingSetExpr) and the other one is (argumentsOfGroupingScalarFunction, argumentsOfAggregateFunction). and use the flattenGroupingSetExpr rewrite all LogicalRepeat outputExpressions, and use the argumentsOfGroupingScalarFunction, argumentsOfAggregateFunction to rewrite only the agg function arguments and the grouping scalar function. So, in the above sql, the pk + 1 AS col_alias3 will not be rewritten to slot, and can be computed.	2024-05-10 22:03:31 +08:00
morrySnow	c0cca6103b	[WIP](test) remove enable_nereids_planner in regression cases (part 2) (#34490 )	2024-05-10 22:02:32 +08:00
kkop	97bc367611	[enhancement](regression-test) modify a key type from BIGINT/LARGEINT to other type (#34436 ) Co-authored-by: cjj2010 <2449402815@qq.com>	2024-05-10 14:48:52 +08:00
zhangstar333	25ae7cd65f	[bug](ipv6) the ipv6 type should be uint128_t (#34121 ) the ipv6 type should be uint128_t, and max value is ffff:ffff:ffff:ffff:ffff:ffff:ffff:ffff if use int128_t type, it's will be min value.	2024-05-10 14:43:46 +08:00
yangshijie	9b712b03b4	[FIX]fix is_ip_address_in_range func with const param (#34266 )	2024-05-10 14:37:20 +08:00
zhangstar333	520774a24b	[fix](serde) fix ipv4/v6 serde functions for arrow, orc, parquet format (#34042 ) this PR is from @sjyango work in #32326, wants merge #32326 into master branch, but it's draft and not maintain long time. so have this new PR. Co-authored-by: sjyango <sjyang2022@zju.edu.cn>	2024-05-10 14:37:04 +08:00
zzzxl	cc00666be6	[opt](inverted index) add inlist condition handling to compound (#34134 ) 1. Previously, the compound did not support the inlist condition, which could impact performance if an inverted index was created.	2024-05-10 14:35:47 +08:00
zy-kkk	5a3107442a	[feature](tvf) support query table value function (#34516 ) (#34640 ) This PR supports a Table Value Function called `Query`. He can push a query directly to the catalog source for execution by specifying `catalog` and `query` without parsing by Doris. Doris only receives the results returned by the query. Currently only JDBC Catalog is supported. Example: ``` Doris > desc function query('catalog' = 'mysql','query' = 'select count() as cnt from test.test'); +-------+--------+------+------+---------+-------+ \| Field \| Type \| Null \| Key \| Default \| Extra \| +-------+--------+------+------+---------+-------+ \| cnt \| BIGINT \| Yes \| true \| NULL \| NONE \| +-------+--------+------+------+---------+-------+ Doris > select from query('catalog' = 'mysql','query' = 'select count(*) as cnt from test.test'); +----------+ \| cnt \| +----------+ \| 30000000 \| +----------+ ```	2024-05-10 14:29:17 +08:00
Mingyu Chen	3ae3f9d6e1	[opt](catalog) support using loading cache for db/table list in external catalog (#33610 ) (#34596 ) bp #33610	2024-05-09 17:50:39 +08:00
wangbo	39fdc9ba0c	[refactor](executor)Rename workload schedule policy #34497	2024-05-08 08:35:20 +08:00
zzzxl	ac56255f82	[opt](inverted index) the "unicode" tokenizer can be configured to disable stop words. (#34467 )	2024-05-07 18:23:43 +08:00
xueweizhang	a33715bc1c	[fix](partial update) only unique table with MOW insert with target columns can consider be a partial update (#33656 ) * [fix](partial update) only unique table with MOW insert with target columns can consider be a partial update Signed-off-by: nextdreamblue <zxw520blue1@163.com> * fix 1 Signed-off-by: nextdreamblue <zxw520blue1@163.com> --------- Signed-off-by: nextdreamblue <zxw520blue1@163.com>	2024-05-07 07:53:25 +08:00
seawinde	92dc8ed718	[opt](mtmv) Add enable materialized view nest rewrite switch (#34197 ) * [opt](mtmv) Add enable materialized view nest rewrite switch * fix ut * fix ut2	2024-05-07 07:51:18 +08:00
feiniaofeiafei	e840102e99	[Feat](nereids)nereids support create table like (#34025 ) nereids support create table like statement. e.g. CREATE TABLE test1.table2 LIKE test1.table1	2024-05-07 07:50:19 +08:00
feiniaofeiafei	fe7d2b8159	[Fix](nereids) ignore slot implements SlotNotFromChildren when check the slot from children in NormalizeAggregate (#34171 )	2024-05-07 07:48:05 +08:00
Pxl	e66dcd0e72	[Bug](materialized-view) change nvl to ifnull when create mv (#34272 ) change nvl to ifnull when create mv	2024-05-07 07:45:33 +08:00
Chester	f7900b53ce	[enhancement](function) floor/ceil/round/round_bankers can use column as scale argument (#34391 )	2024-05-06 22:18:36 +08:00
Jerry Hu	7248420cfd	[chore](session_variable) Add 'data_queue_max_blocks' to prevent the DataQueue from occupying too much memory. (#34017 ) (#34395 )	2024-05-05 21:20:33 +08:00
Mingyu Chen	8da260ee0d	[fix](hdfs)read 'fs.defaultFS' from core-site.xml for hdfs load which has no default fs (#34217 ) (#34372 ) bp #34217 Co-authored-by: slothever <18522955+wsjz@users.noreply.github.com>	2024-05-01 00:31:49 +08:00
Mingyu Chen	35f8563a75	[feature](iceberg) support iceberg equality delete (#34223 ) (#34327 ) bp #34223 Co-authored-by: Ashin Gau <AshinGau@users.noreply.github.com>	2024-04-30 11:51:29 +08:00
Qi Chen	7cb00a8e54	[Feature](hive-writer) Implements s3 file committer. (#34307 ) Backport #33937.	2024-04-29 19:56:49 +08:00
daidai	1bfe0f0393	[feature](iceberg)support read iceberg complex type，iceberg.orc format and position delete. (#33935 ) (#34256 ) master #33935	2024-04-29 14:40:12 +08:00
amory	20bd0c2987	[FIX](cases )fix ipv6 value for regress case	2024-04-29 13:37:29 +08:00
Qi Chen	99af54f779	[Fix](orc-reader) Fix the issue when string col has mixed plain and dict encoding in different stripes. (#34146 ) (#34248 ) backport #34146	2024-04-28 19:43:57 +08:00
苏小刚	11039ade7b	[opt](paimon) support mapping Paimon column type "Row" to Doris type "Struct" (#34239 ) backport: #33786	2024-04-28 19:38:50 +08:00
zy-kkk	1fda68f738	[feature](planner) Support `select constant from dual` syntax sugar (#34200 ) (#34232 ) In MySQL, it's common to use a simplified syntax like `SELECT constant FROM dual` which is equivalent to just `SELECT constant`. This syntax is often used by BI tools when utilizing MySQL connectors to verify connection validity. To enhance compatibility and ensure seamless integration with such tools, we have now implemented this feature in Doris. ### Key Changes: - Doris now interprets `SELECT constant FROM dual` as `SELECT constant`, aligning with MySQL's behavior. - This update ensures that BI tools can use standard MySQL connectors without modifications or errors when connecting to Doris.	2024-04-28 15:56:16 +08:00
Mingyu Chen	45556686ea	[fix](test) fix some external test cases (#34209 ) Fix some test cases and enable `test_information_schema_external` suite	2024-04-27 23:25:33 +08:00
Xin Liao	cd1c9edd71	[fix](pipeline-load) fix no error url when data quality error and total rows is negative (#34072 ) (#34204 ) Co-authored-by: HHoflittlefish777 <77738092+HHoflittlefish777@users.noreply.github.com>	2024-04-27 18:19:08 +08:00
Luwei	36e80af327	[fix](schema change) fix the defineName field is not the same when copying column (#34201 ) * [fix](schema change) fix the defineName field is not the same when copying column * fix	2024-04-27 11:59:07 +08:00
qiye	414fbd353e	[fix](ES catalog)Make col != '' behavior consistent with SQL (#34151 ) In SQL syntax, `col != ''` equals `col.length() > 0`. It means that this column must exist in ES doc fields and its content is not empty. In this PR, we make a special translation for this binary predicate to keep the behavior of both consistent. --------- Co-authored-by: Luennng <luennng@gmail.com>	2024-04-27 02:29:33 +08:00
xzj7019	c125148deb	[opt](Nereids) bucket shuffle downgrade expansion (#34088 ) Expand bucket shuffle downgrade condition, which originally requiring a single partition after pruning, basic table and bucket number < para number. Currently, we expect this option can be used for disabling bucket shuffle more efficiently, without above restrictions. Co-authored-by: zhongjian.xzj <zhongjian.xzj@zhongjianxzjdeMacBook-Pro.local>	2024-04-27 02:29:33 +08:00
苏小刚	0f0c0a266b	[opt](parquet)Skip page with offset index (#33082 ) Make skip_page() in ColumnChunkReader more efficient. No more reading page headers if there are pagelocations in chunk.	2024-04-26 15:06:16 +08:00
Qi Chen	acc2b532e7	[Test](hive-writer) Adjust test_hive_write_partitions regression test to resolve special characters issue with git on windows. (#34026 )	2024-04-26 15:05:47 +08:00
morrySnow	b24ff9953d	[fix](Nereids) column pruning should prune map in cte consumer (#34079 ) we save bi-map in cte consumer to get the maping between producer and consumer. the consumer's output is decided by the map in it. so, cte consumer should be output prunable, and should remove useless entry from map when do column pruning	2024-04-26 15:05:19 +08:00
feiniaofeiafei	b41a5339d3	[Fix](nereids) fix rule merge_aggregate when has project (#33892 )	2024-04-26 15:05:09 +08:00
Mingyu Chen	50f9d47e96	[test](hive) run suite cases both in hive2 and hive3 (#33874 ) (#34156 ) bp #33874 Co-authored-by: 苏小刚 <suxiaogang223@icloud.com>	2024-04-26 13:48:09 +08:00

1 2 3 4 5 ...

2879 Commits