doris

Author	SHA1	Message	Date
airborne12	0cd5183556	[Refactor](inverted index) refact tokenize function for inverted index (#22313 )	2023-08-02 19:12:22 +08:00
Mryange	ddd90855a9	[vectorized](udaf) java udaf support with map type (#22397 ) [vectorized](udaf) java udaf support with map type (#22397) * test * remove some unused * update * add case	2023-08-02 15:03:44 +08:00
Mryange	bf50f9fa7f	[fix](decimal) fix cast rounding half up with negative number (#22450 )	2023-08-01 21:47:42 +08:00
qiye	b8399148ef	[fix](DOE) es catalog not working with pipeline,datetimev2, array and esquery (#22046 )	2023-08-01 21:45:16 +08:00
minghong	d5d82b7c31	[stats](nereids) fix bug for avg-size (#22421 )	2023-08-01 17:13:00 +08:00
Pxl	8d16f1bb09	[Chore](materialized-view) update documentation about materialized-view and update test (#22350 ) update documentation about materialized-view and update test	2023-08-01 15:13:34 +08:00
Gabriel	7a2ff56863	[regression](fix) fix test_round case (#22441 )	2023-08-01 11:35:44 +08:00
Jerry Hu	c1f36639fd	[fix](sort) VSortedRunMerger does not return any rows with a large offset value (#22191 )	2023-07-31 22:28:13 +08:00
starocean999	450e0b1078	[fix](nereids) recompute logical properties in plan post process (#22356 ) join commute rule will swap the left and right child. This cause the change of logical properties. So we need recompute the logical properties in plan post process to get the correct result	2023-07-31 21:04:39 +08:00
LiBinfeng	3a1d678ca9	[Fix](Planner) fix parse error of view with group_concat order by (#22196 ) Problem: When create view with projection group_concat(xxx, xxx order by orderkey). It will failed during second parse of inline view For example: it works when doing "SELECT id, group_concat(`name`, "," ORDER BY id) AS test_group_column FROM test GROUP BY id" but when create view it does not work "create view test_view as SELECT id, group_concat(`name`, "," ORDER BY id) AS test_group_column FROM test GROUP BY id" Reason: when creating view, we will doing parse again of view.toSql() to check whether it has some syntax error. And when doing toSql() to group_concat with order by, it add seperate ', ' between second parameter and order by. So when parsing again, it would failed because it is different semantic with original statement. group_concat(`name`, "," ORDER BY id) ==> group_concat(`name`, "," , ORDER BY id) Solved: Change toSql of group_concat and add order by statement analyze() of group_concat in Planner cause it would work if we get order by from view statement and do not analyze and binding slot reference to it	2023-07-31 17:20:23 +08:00
amory	7261845b3d	[FIX](complex-type)fix complex type nested col_const (#22375 ) for array/map/struct in mysql_writer unpack_if_const only unpack self column not nested , so col_const should not used in nested column.	2023-07-31 14:53:18 +08:00
zclllyybb	f2919567df	[feature](datetime) Support timezone when insert datetime value (#21898 )	2023-07-31 13:08:28 +08:00
TengJianPing	79289e32dc	[fix](cast) fix wrong result of casting empty string to array date (#22281 )	2023-07-30 21:15:03 +08:00
Jibing-Li	03761c37cd	[Improvement](multi catalog) Support Iceberg, Paimon and MaxCompute table in nereids. (#22338 )	2023-07-29 21:43:35 +08:00
Mryange	47c2cc5c74	[vectorized](udf) java udf support with return map type (#22300 )	2023-07-29 12:52:27 +08:00
daidai	ae8a26335c	[opt](hive)opt select count() stmt push down agg on parquet in hive . (#22115 ) Optimization "select count() from table" stmtement , push down "count" type to BE. support file type : parquet ，orc in hive . 1. 4kfiles , 60kwline num before: 1 min 37.70 sec after: 50.18 sec 2. 50files , 60kwline num before: 1.12 sec after: 0.82 sec	2023-07-29 00:31:01 +08:00
zhannngchen	53d255f482	[fix](partial update) remove CHECK on illegal number of partial columns (#22319 )	2023-07-28 23:11:58 +08:00
xzj7019	f7c106c709	[opt](nereids) enhance broadcast join cost calculation (#22092 ) Enhance broadcast join cost calculation, by considering both the build side effort from building bigger hash table, and more probe side effort from bigger cost of ProbeWhenBuildSideOutput and ProbeWhenSearchHashTable, if parallel_fragment_exec_instance_num is more than 1. Current solution gives a penalty factor on rightRowCount, and the factor is the total instance number to the power of 2. Penalty on outputRows is not taken currently and will be refined in next generation cost model. Also brings some update for shape checking: update original control variable in shape file parallel_fragment_exec_instance_num to parallel_pipeline_task_num, if pipeline is enabled. fix a be_number variable inactive issue.	2023-07-28 23:06:02 +08:00
Kaijie Chen	2f43e59535	[test](regression) add partial update seq_col delete cases (#22340 )	2023-07-28 17:36:55 +08:00
HHoflittlefish777	05abfbc5ef	[improvement](regression-test) add compression algorithm regression test (#22303 )	2023-07-28 17:28:52 +08:00
starocean999	5a0ad09856	[fix](nereids) SubqueryToApply may lost conjunct (#22262 ) consider sql: ``` SELECT * FROM sub_query_correlated_subquery1 t1 WHERE coalesce(bitand( cast( (SELECT sum(k1) FROM sub_query_correlated_subquery3 ) AS int), cast(t1.k1 AS int)), coalesce(t1.k1, t1.k2)) is NULL ORDER BY t1.k1, t1.k2; ``` is Null conjunct is lost in SubqueryToApply rule. This pr fix it	2023-07-28 15:08:56 +08:00
bobhan1	0c734a861e	[Enhancement](delete) eliminate reading the old values of non-key columns for delete stmt (#22270 )	2023-07-28 14:37:33 +08:00
morrySnow	5da5fac37a	[refactor](Nereids) add result sink node (#22254 ) use ResultSink as query root node to let plan of query statement has the same pattern with insert statement	2023-07-28 11:31:09 +08:00
zhangy5	adc44d9f46	[regression-test] add list partition case and multi partition keys case (#22042 ) * [regression-test] add list partition case and multi partition keys case * fix delete failed	2023-07-28 10:12:35 +08:00
Ashin Gau	0d7d9b92db	[fix](multi-catalog) complex types parsing failed, with unexpected nulls and rows (#22228 ) Fix tow bugs: 1. Unexpected null values in array column. If 65535 consecutive values are not null in nullable array column, this error will be triggered. The reason is that the array parser did not handle boundary conditions. 2. The number of rows of key filed, and that of value field in map column are not equal. Similarly, the number of rows among fields in struct column are not the same. This would be triggered when the number of rows are not equal among parquet pages of different columns in a row group.	2023-07-28 10:03:08 +08:00
Qi Chen	8caa5a9ba4	[Fix](mutli-catalog) Fix null partitions error in iceberg tables. (#22185 ) ### Issue when partition has null partitions, it throws error `Failed to fill partition column: t_int=null` ### Resolution - Fix the following null partitions error in iceberg tables by replacing null partition to '\N'. - Add regression test for hive null partition.	2023-07-27 23:57:35 +08:00
Jerry Hu	b5fa29e138	[fix](bitmap) incorrect result of function 'bitmap_from_array' (#22305 )	2023-07-27 22:44:06 +08:00
谢健	716d58f5ff	[fix](Nereids) decimal divide should not return null if numerator is zero (#22309 )	2023-07-27 20:23:04 +08:00
Jibing-Li	a87d34b19b	[Fix](multi catalog statistics)Improve external table statistics collection (#22224 ) Improve external table statistics collection, including log, observability and fix some bugs. 1. Add Running state for statistics job. 2. Add progress for show analyze job. (n/m tasks finished, n/m task failed and so on) 3. Add analyze time cost for show analyze task. 4. Make task failure message more clear. 5. Synchronize the job status updating code in updateTaskStatus. 6. Fix NPE in HMSAnalyzeTask. (Avoid refreshing statistics cache if the collection sql failed) 7. Return error message for with sync collection while timeout. 8. Log level improvement 9. Fix misuse of logCreateAnalysisJob for tasks.	2023-07-27 20:01:14 +08:00
morrySnow	ae5e39ad26	[opt](Nereids) add double signature back for round like function (#22284 ) add double signature back for round like function	2023-07-27 19:10:43 +08:00
lsy3993	6f1c03c766	[fix](jdbc_catalog) fix int and bigint in mysql view when use doris catalog (#22251 )	2023-07-27 16:50:42 +08:00
Kaijie Chen	0512e0b168	[test](regression) add cases for partial update with sequence_type (#22215 )	2023-07-27 15:51:01 +08:00
lsy3993	4f6a3c5bf0	[feature](catalog) support clob type in oracle jdbc catalog (#21532 )	2023-07-27 15:49:15 +08:00
zhangstar333	ddfdf62993	[opt](planner) support to parse scientific notation(aEb) (#22248 )	2023-07-27 13:31:37 +08:00
wuwenchi	41a230b721	[fix] iceberg catalog to specify the version and time (#22209 ) problem: 1. create a iceberg_type catalog: 2. use iceberg catalog to specify verison ``` mysql> show catalog iceberg; +----------------------+--------------------------+ \| Key \| Value \| +----------------------+--------------------------+ \| type \| iceberg \| \| iceberg.catalog.type \| hms \| \| hive.metastore.uris \| thrift://127.0.0.1:9083 \| \| hadoop.username \| hadoop \| \| create_time \| 2023-07-25 16:51:00.522 \| +----------------------+--------------------------+ 5 rows in set (0.02 sec) mysql> select * from iceberg.iceberg_db.tb1 FOR VERSION AS OF 8783036402036752909; ERROR 5090 (42000): errCode = 2, detailMessage = Only iceberg/hudi external table supports time travel in current version ``` change: Add `ICEBERG_EXTERNAL_TABLE` type for specify the version and time	2023-07-27 12:04:41 +08:00
zy-kkk	619a2857e1	[improvement](jdbc catalog) improve mysql jdbc catalog read bytea`s types & else improve (#22233 )	2023-07-27 10:18:37 +08:00
Gabriel	341c45974c	[round](decimalv2) round precise decimalv2 value (#22258 )	2023-07-27 10:00:36 +08:00
Xinyi Zou	163a38a527	[opt](Nereids) support sql cache (#22144 ) 1. let Nereids support sql cache 2. let legacy planner's sql cache supports union all	2023-07-27 09:57:31 +08:00
Siyang Tang	8fb28ecc9e	[test](partial-update) add some cases for partial-update (#22210 )	2023-07-27 09:52:40 +08:00
HHoflittlefish777	dcd6844ea5	[improvement](regression-test) add partial update with schema change case (#22213 )	2023-07-27 09:51:42 +08:00
zhangstar333	fb41265c27	[opt](Nereids) add boolean type signature for sum aggregate function (#21959 )	2023-07-27 09:41:19 +08:00
TengJianPing	8ff487cc4b	[fix](cast) fix invalid value error when casting null date value to string then casting to date value (#22223 )	2023-07-26 17:59:01 +08:00
morrySnow	14dcc53135	[fix](Nereids) cast time should turn nullable on all valid types (#22242 ) valid types to cast to time/timev2: - TINYINT - SMALLINT - INT - BIGINT - LARGEINT - FLOAT - DOUBLE - CHAR - VARCHAR - STRING	2023-07-26 17:56:19 +08:00
bobhan1	be69025878	[opt](Nereids) add partial update support for delete stmt (#22184 ) Currently, the new optimizer don't consider anything about partial update. This PR add the ability to convert a delete statement to a partial update insert statement for merge-on-write unique table	2023-07-26 17:34:31 +08:00
jakevin	bb67a1467a	[fix](Nereids): mergeGroup should merge target Group into existed Group (#22123 )	2023-07-26 13:13:25 +08:00
morrySnow	21a3593a9a	[fix](Nereids) translate failed when enable topn two phase opt (#22197 ) 1. should not add rowid slot to reslovedTupleExprs 2. should set notMaterialize to sort's tuple when do two phase opt	2023-07-26 11:38:50 +08:00
zy-kkk	cf677b327b	[fix](jdbc catalog) Fixed mappings with type errors for bool and tinyint(1) (#22089 ) First of all, mysql does not have a boolean type, its boolean type is actually tinyint(1), in the previous logic, We force tinyint(1) to be a boolean by passing tinyInt1isBit=true, which causes an error if tinyint(1) is not a 0 or 1, Therefore, we need to match tinyint(1) according to tinyint instead of boolean, and this change will not affect the correctness of where k = 1 or where k = true queries	2023-07-25 22:45:22 +08:00
zhengyu	5c8eda8685	[enhencement](regression) add UPDATE & DELETE tests for MOW partial update (#22212 )	2023-07-25 22:03:38 +08:00
airborne12	fc2b9db0ad	[Feature](inverted index) add tokenize function for inverted index (#21813 ) In this PR, we introduce TOKENIZE function for inverted index, it is used as following: ``` SELECT TOKENIZE('I love my country', 'english'); ``` It has two arguments, first is text which has to be tokenized, the second is parser type which can be english, chinese or unicode. It also can be used with existing table, like this: ``` mysql> SELECT TOKENIZE(c,"chinese") FROM chinese_analyzer_test; +---------------------------------------+ \| tokenize(`c`, 'chinese') \| +---------------------------------------+ \| ["来到", "北京", "清华大学"] \| \| ["我爱你", "中国"] \| \| ["人民", "得到", "更", "实惠"] \| +---------------------------------------+ ```	2023-07-25 15:05:35 +08:00
mch_ucchi	d96e31c4d7	[opt](Nereids) not push down global limit to avoid early gather (#21891 ) the global limit will create a gather action, and all the data will be calculated in one instance. If we push down the global limit, the node run after the limit node will run slowly. We fix it by push down only local limit. a join plan tree before fixing: ``` LogicalLimit(global) LogicalLimit(local) Plan() LogicalLimit(global) LogicalLimit(local) LogicalJoin LogicalLimit(global) LogicalLimit(local) Plan() LogicalLimit(global) LogicalLimit(local) Plan() after fixing: LogicalLimit(global) LogicalLimit(local) Plan() LogicalLimit(local) LogicalJoin LogicalLimit(local) Plan() LogicalLimit(local) Plan() ```	2023-07-25 14:45:20 +08:00

1 2 3 4 5 ...

1494 Commits