doris

Author	SHA1	Message	Date
zzzzzzzs	cbbad5d95c	[typo](doc)Update SHOW-PROC.md and SHOW-CATALOGS.md (#18398 )	2023-04-05 22:24:35 +08:00
Ashin Gau	47aa8a6d8a	[fix](file_cache) turn on file cache by FE session variable (#18340 ) Fix tow bugs: 1. Enabling file caching requires both `FE session` and `BE` configurations(enable_file_cache=true) to be enabled. 2. `ParquetReader` has not used `IOContext` previously, but `CachedRemoteFileReader::read_at` needs `IOContext` after PR(#17586).	2023-04-05 15:51:47 +08:00
Pxl	0a4381197a	[Bug](MTMV) fix waitingMTMVTaskFinished failed at test_mtmv_ssb_ddl (#18373 ) fix waitingMTMVTaskFinished failed at test_mtmv_ssb_ddl	2023-04-05 11:04:41 +08:00
abmdocrt	1ec400c786	[fix](SSL) fix ssl connection buffer overflow (#18359 )	2023-04-05 08:42:41 +08:00
HOHO	668031986b	Update install-faq.md (#18385 )	2023-04-05 08:38:06 +08:00
Sebastian Gassner	4edf2acc81	[typo](doc)Fixing broken links in docker cluster docs, improving formatting. (#18290 )	2023-04-05 08:36:47 +08:00
Jibing-Li	ea60d65384	[Improvement](multi catalog)Move split size config to session variable (#18355 ) Move split size config to session variable. Before, it was in Config class, user need to restart FE after change it.	2023-04-05 01:02:47 +08:00
gitccl	7f8d92656e	[fix](streamload) fix stream load failed when enable profile (#18364 ) #18015 enables stream load profile log, however be will encounter rpc fail when loading tpch data(see #18291). This is because when `is_report_success` is true, be will reportExecStatus to fe, but fe cannot find QueryInfo in `coordinatorMap`, thus it will return error to be.	2023-04-05 01:01:46 +08:00
xueweizhang	d8b293de07	[fix](multi-catalog) add catalog info for show proc (#18276 ) Signed-off-by: nextdreamblue <zxw520blue1@163.com>	2023-04-04 22:49:22 +08:00
huangzhaowei	7c36bef6bc	[Feature-Wip](MySQL Load)Show load warning for my sql load (#18224 ) 1. Support the show load warnings for mysql load to get the detail error message. 2. Fix fillByteBufferAsync not mark the load as finished in same data load 3. Fix drain data only in client mode.	2023-04-04 22:44:48 +08:00
morrySnow	e29fc3b46b	[fix](chore) fix compile failed in JdbcExecutor and revert #18306 since be crash randomly (#18371 ) fix 2 problems: 1. PR #18187 use the api resizeColumn in JNINativeMethod has been removed by #17960 2. revert PR #18306 to fix pipeline core when load	2023-04-04 20:04:28 +08:00
Ashin Gau	66bfd18601	[opt](file_reader) add prefetch buffer to read csv&json file (#18301 ) Co-authored-by: ByteYue <[yj976240184@gmail.com](mailto:yj976240184@gmail.com)> This PR is an optimization for https://github.com/apache/doris/pull/17478: 1. Change the buffer size of `LineReader` to 4MB to align with the size of prefetch buffer. 2. Lazily prefetch data in the first read to prevent wasted reading. 3. S3 block size is 32MB only, which is too small for a file split. Set 128MB as default file split size. 4. Add `_end_offset` for prefetch buffer to prevent wasted reading. The query performance of reading data on object storage is improved by more than 3x+.	2023-04-04 19:05:22 +08:00
zclllyybb	d7623028e9	[doc](developer-guide) add some debug tricks to dev-guide (#18225 ) add method to debug core-dump file in vscode. and some BE debug tricks.	2023-04-04 17:10:34 +08:00
minghong	3fc8c19735	[improve](nereids)compute statsRange.length() according to the column datatype (#18331 ) we map date/datetime/V2 to double. this map reserves date order, but it does not reserve range length. For example, from 1990-01-01 to 1991-01-01, there are 12 months. for filter `A < 1990-02-01`, the selectivity should be `1/12`. if we compute this filter by their corresponding double value, `sel = (19900201 - 19900101) / (19910101 - 19900101) = 100/10000 = 1/100` the error is about 10 times. This pr aims to fix this error. Describe your changes. Solution: convert double to its corresponding dataType(date/datev2), then compute the range length with respect to its datatype.	2023-04-04 14:20:34 +08:00
zhannngchen	175e5d405c	[improvement](merge-on-write) remove CHECK if lookup_row_key return unexpected status (#18326 )	2023-04-04 12:42:07 +08:00
yixiutt	87e83081ff	[test](compaction) add delete test (#18335 )	2023-04-04 12:28:19 +08:00
yixiutt	0cada3f81d	[Enhancement](compaction) return error instead of core when ctx not valid (#18363 )	2023-04-04 12:27:13 +08:00
zhangstar333	54dbb4af67	[vectorzied](jdbc) refactor jdbc table read array type (#18187 ) jdbc read array type get result from Doris is string, PG is java.sql.array, CK is java.lang.object it's difficult to maintain and read the code, so change all database's array result to string, then add a cast function from string to doris array type	2023-04-04 11:57:04 +08:00
Xin Liao	418ea0a24e	[fix](merge-on-write) fix that failed to capture_consistent_rowsets when full clone (#18346 ) When full clone, if the max version of the local table is less than or equal to the max version of the clone table, there is no need to calculate the delete bitmap again.	2023-04-04 10:39:28 +08:00
Mingyu Chen	2a301eb437	[deps](arrow) update arrow download link (#18360 )	2023-04-04 10:39:04 +08:00
zhangstar333	50e6c4216a	[vectorized](function) suppoort date_trunc function truncate week mode (#18334 ) support date_trunc could truncate week eg: select date_trunc('2023-4-3 19:28:30', 'week');	2023-04-04 10:24:26 +08:00
Gabriel	a724443eb9	[Improvement](predicate) optimize short-circuit predicates (#18278 ) For scan node with no vectorized predicate, the input column for the first short-circuit predicate is dense and we don't need to access the selector column. This PR improve performance by ~30% on TPCH Q3.	2023-04-04 10:21:41 +08:00
yongkang.zhong	6231ca80f7	[improve](clickhouse catalog) Add `"` wrap select column for the sql query clickhouse jdbc (#18352 )	2023-04-04 10:19:24 +08:00
Lightman	af80e65094	[Improve](FileCahe) Support the file cache profile in olap scan node and Update the profile (#17710 ) We want to use file cache for caching cold data in S3. When reading them, we want to know where the data come from and the time taken to read the datas. So we support the metrics in olap scan node. And for clearing the information, i also update the fields about the metrics.	2023-04-04 10:18:30 +08:00
minghong	3e7a9424e4	[feature](nereids) explain shape plan (#18296 ) `explain shape plan select ...` only print plan shape related information, including - node name - join type, join condition - filter condition - agg phase It is painful to maintain regression cases using explain since there are a lot of mutable information, like slot id. By this pr, we could use explain shape plan in regression cases. for exmaple: this is tpch q2 +-----------------------------------------------------------------------------------------------------------+ \| Explain String \| +-----------------------------------------------------------------------------------------------------------+ \| PhysicalTopN \| \| --PhysicalDistribute \| \| ----PhysicalTopN \| \| ------PhysicalProject \| \| --------filter((cast(ps_supplycost as DECIMAL(27, 9)) = min(ps_supplycost) OVER(PARTITION BY p_partkey))) \| \| ----------PhysicalWindow \| \| ------------PhysicalQuickSort \| \| --------------PhysicalProject \| \| ----------------hashJoin[INNER_JOIN](supplier.s_suppkey = partsupp.ps_suppkey) \| \| ------------------PhysicalProject \| \| --------------------hashJoin[INNER_JOIN](part.p_partkey = partsupp.ps_partkey) \| \| ----------------------PhysicalProject \| \| ------------------------PhysicalOlapScan[partsupp] \| \| ----------------------PhysicalProject \| \| ------------------------filter((part.p_size = 15)(p_type like '%BRASS')) \| \| --------------------------PhysicalOlapScan[part] \| \| ------------------PhysicalDistribute \| \| --------------------hashJoin[INNER_JOIN](supplier.s_nationkey = nation.n_nationkey) \| \| ----------------------PhysicalOlapScan[supplier] \| \| ----------------------PhysicalDistribute \| \| ------------------------hashJoin[INNER_JOIN](nation.n_regionkey = region.r_regionkey) \| \| --------------------------PhysicalProject \| \| ----------------------------PhysicalOlapScan[nation] \| \| --------------------------PhysicalDistribute \| \| ----------------------------PhysicalProject \| \| ------------------------------filter((region.r_name = 'EUROPE')) \| \| --------------------------------PhysicalOlapScan[region] \| +-----------------------------------------------------------------------------------------------------------+	2023-04-04 09:44:15 +08:00
xueweizhang	798d2e5160	[fix](catalog) all properties should be checked when create unpartitioned table (#18149 ) all properties should be checked when create unpartitioned table like partitioned table. Signed-off-by: nextdreamblue <zxw520blue1@163.com>	2023-04-04 08:53:45 +08:00
ZhangYu0123	8b85c55117	[vectorized](function) Support array_shuffle and shuffle function. (#18116 ) --------- Co-authored-by: zhangyu209 <zhangyu209@meituan.com>	2023-04-04 08:53:13 +08:00
Qi Chen	eb0fd0017e	[Fix](orc-reader) Fix the scale of decimal column is incorrect when query orc tables. (#18324 ) The scale of decimal column is incorrect when query orc tables.	2023-04-04 08:50:47 +08:00
wangbo	fc407f4afe	[improvement](executor) Reduce ScannnerCtx Scheduling times (#18306 ) * remove sche in scan operator	2023-04-03 22:54:34 +08:00
starocean999	88c5e64c4a	[fix](nereids) fix bug of SelectMaterializedIndexWithAggregate rule (#18265 ) 1. create a project node to adjust the output column position when a mv is selected in olap scan node 2. pass SlotReference's column info when call Alias's toSlot() method 3. should compare plan's logical properties when compare two plans after rewrite	2023-04-03 22:32:43 +08:00
Jerry Hu	1e51af0784	[fix](scan) Avoid using incorrect cache code in ComparisonPredicate (#18332 ) * [fix](scan) Avoid using incorrect cache code in ComparisonPredicate * recovery the regression test	2023-04-03 20:37:35 +08:00
Xinyi Zou	dd78001cc1	[fix](memory) Fix memtable flush mem tracker #18330	2023-04-03 20:37:14 +08:00
yongkang.zhong	fe9d2b00fc	[test](jdbc catalog) add clickhouse jdbc catalog base type test (#18007 )	2023-04-03 20:18:36 +08:00
ZhangYu0123	b627088e8c	[Optimization](String) Optimize q20 q21 q22 q23 LIKE_SUBSTRING (like '%xxx%') (#18309 ) Optimize q20, q21, q22, q23 LIKE_SUBSTRING (like '%xxxx%'). Idea is from clickhouse stringsearcher: Stringsearcher is about 10%~20% faster than volnitsky algorithm when needle size is less than 10 using two chars at beginning search in SIMD . Stringsearcher is faster than volnitsky algorithm, when needle size is less than 21. The changes are as follows: Using first two chars of needle at beginning search. We can compare two chars of needle and [n:n+17) chars in haystack in SIMD in one loop. Filter efficiency will be higher. When env support SIMD, we use stringsearcher. Test result in clickbench: q20 is about 15% up. q20: SELECT COUNT() FROM hits WHERE URL LIKE '%google%'; q21, q22 is about 1%~5% up. q21: SELECT SearchPhrase, MIN(URL), COUNT() AS c FROM hits WHERE URL LIKE '%google%' AND SearchPhrase <> '' GROUP BY SearchPhrase ORDER BY c DESC LIMIT 10; q22: SELECT SearchPhrase, MIN(URL), MIN(Title), COUNT() AS c, COUNT(DISTINCT UserID) FROM hits WHERE Title LIKE '%Google%' AND URL NOT LIKE '%.google.%' AND SearchPhrase <> '' GROUP BY SearchPhrase ORDER BY c DESC LIMIT 10; q23 is about 30%~40% up and not stable. q23: SELECT FROM hits WHERE URL LIKE '%google%' ORDER BY EventTime LIMIT 10;	2023-04-03 18:09:15 +08:00
yongkang.zhong	eb6dbc03e0	[typo](docs) add regression test doc & fix api doc (#18329 )	2023-04-03 17:40:41 +08:00
ZhangYu0123	d4688620e9	[opt](array) optimize array_sortby using qsort instead of bubble sort #18311	2023-04-03 17:10:51 +08:00
Gabriel	96a64dc9e8	[Improvement](pipeline) Use bloom runtime filter by default for pipeline engine (#18177 )	2023-04-03 15:31:48 +08:00
Gabriel	368a2f7ace	[Bug](decimal) Fix string to decimal (#18282 )	2023-04-03 15:30:48 +08:00
caoliang-web	3078ee1854	[regression](decimalv3)Add decimal type as filter condition in regression test (#17160 ) Add decimal type as filter condition in regression test	2023-04-03 14:20:09 +08:00
yongjinhou	aff260c06f	[Enhancement](HttpServer) Support https interface (#16834 ) 1. Organize http documents 2. Add http interface authentication for FE 3. Support https interface for FE 4. Provide authentication interface 5. Add http interface authentication for BE 6. Support https interface for BE	2023-04-03 14:18:17 +08:00
Mingyu Chen	ecd3fd07f6	[feature](colocate) support cross database colocate join (#18152 )	2023-04-03 14:03:42 +08:00
Jibing-Li	e260dca7a1	[Improvement](multi catalog)Change hive metastore cache split value type to Doris defined Split. Fix split file length -1 bug (#18319 ) HiveMetastoreCache type for file split was Hadoop InputSplit. In this pr, change it to Doris defined Split This change could avoid convert it every time. Also fix the explain verbose result return -1 for split file length.	2023-04-03 13:54:28 +08:00
Xin Liao	6677841b7e	[fix](merge-on-write) fix that failed to capture_consistent_rowsets when revise tablet meta (#18283 ) Should modify _timestamped_version_tracker firstly before capture_consistent_rowsets when update delete bitmap in revise_tablet_meta.	2023-04-03 13:02:34 +08:00
Liqf	961f5d1bb7	[feature](function)Add St_Angle/St_Azimuth function (#18293 ) Add St_Angle/St_azimuth function： St_Angle： Enter three point, which represent two intersecting lines. Returns the angle between these lines. Point 2 and point 1 represent the first line and point 2 and point 3 represent the second line. The angle between these lines is in radians, in the range [0, 2pi). The angle is measured clockwise from the first line to the second line. ` mysql> SELECT ST_Angle(ST_Point(1, 0),ST_Point(0, 0),ST_Point(0, 1)); +----------------------------------------------------------------------+ \| st_angle(st_point(1.0, 0.0), st_point(0.0, 0.0), st_point(0.0, 1.0)) \| +----------------------------------------------------------------------+ \| 4.71238898038469 \| +----------------------------------------------------------------------+ 1 row in set (0.04 sec) ` St_azimuth： Enter two point, and returns the azimuth of the line segment formed by points 1 and 2. The azimuth is the angle in radians measured between the line from point 1 facing true North to the line segment from point 1 to point 2. ` mysql> SELECT st_azimuth(ST_Point(0, 0),ST_Point(1, 0)); +----------------------------------------------------+ \| st_azimuth(st_point(0.0, 0.0), st_point(1.0, 0.0)) \| +----------------------------------------------------+ \| 1.5707963267948966 \| +----------------------------------------------------+ 1 row in set (0.04 sec)	2023-04-03 13:01:59 +08:00
Pxl	e77833bfa1	[Bug](materialized-view) fix where clause persistence replay incorrect (#18228 ) fix where clause persistence replay incorrect	2023-04-03 12:49:01 +08:00
zhangstar333	94e3472050	[bug](function) fix count equal function return incorrect value (#18200 ) fix count equal function return incorrect value	2023-04-03 11:20:36 +08:00
AKIRA	ce4dc681be	[test](stats) Test framework for stats estimation on TPCH-1G dataset (#18267 ) Implement a test framework for stats estimation on TPCH-1G dataset to ensure accuracy	2023-04-03 11:01:57 +08:00
WenYao	2bce4db81a	[Enchancement](mysql-compatable) add regression-test for MySQLdump #18208 add regression-test for like this: mysqldump -h127.0.0.1 -P9030 -uroot --no-tablespaces --databases > /backup/mysqldump/test.db To prevent errors Unknown table 'column_statistics' in information_schema (1109), the table information_schema.column_statistics was added.	2023-04-03 09:49:07 +08:00
TengJianPing	7cd8f7c9ba	[fix](grouping) fix coredump of grouping function for outer join (#18292 ) Result of functions grouping and grouping_id is always not nullable, but outer join will convert the result column to nullable when necessary, which will cause mismatch of column type and column object when executing unctions grouping and grouping_id.	2023-04-03 09:35:31 +08:00
Xin Liao	b66e9f8906	[fix](load) handle null map right in OlapDataConvertor (#18236 ) The offset of _nullmap and _value are inconsistent in OlapDataConvertor, so the obtained null flag is incorrect when calling get_ data_ at function. When the key column or sequence column has null values, the encoding of the short key index or primary key index may be wrong. This was introduced by #10883 #10925.	2023-04-03 09:14:05 +08:00

... 170 171 172 173 174 ...

18263 Commits