doris

Author	SHA1	Message	Date
Jerry Hu	c22d097b59	[improvement](compress) Support compress/decompress block with lz4 (#11955 )	2022-08-22 17:35:43 +08:00
Ashin Gau	6d925054de	[feature-wip](parquet-reader) decode parquet time & datetime & decimal (#11845 ) 1. Spark can set the timestamp precision by the following configuration: spark.sql.parquet.outputTimestampType = INT96(NANOS), TIMESTAMP_MICROS, TIMESTAMP_MILLIS DATETIME V1 only keeps the second precision, DATETIME V2 keeps the microsecond precision. 2. If using DECIMAL V2, the BE saves the value as decimal128, and keeps the precision of decimal as (precision=27, scale=9). DECIMAL V3 can maintain the right precision of decimal	2022-08-22 10:15:35 +08:00
camby	83ea4ea984	[refractor](bitmap) bitmap serialize and deserialize refractor (#11921 ) Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-08-22 08:52:20 +08:00
slothever	124b4f7694	[feature-wip](parquet-reader) row group reader ut finish (#11887 ) Co-authored-by: jinzhe <jinzhe@selectdb.com>	2022-08-18 17:18:14 +08:00
slothever	f39f57636b	[feature-wip](parquet-reader) update column read model and add page index (#11601 )	2022-08-16 15:04:07 +08:00
lihangyu	01383c3217	[Enhancement](stream-load-json) using simdjson to parse json (#11665 ) Currently we use rapidjson to parse json document, It's fast but not fast enough compare to simdjson.And I found that the simdjson has a parsing front-end called simdjson::ondemand which will parse json when accessing fields and could strip the field token from the original document, using this feature we could reduce the cost of string copy(eg. we convert everthing to a string literal in _write_data_to_column by sprintf, I saw a hotspot from the flamegrame in this function, using simdjson::to_json_string will strip the token(a string piece) which is std::string_view and this is exactly we need).And second in _set_column_value we could iterate through the json document by for (auto field: object_val) {xxx}, this is much faster than looking up a field by it's field name like objectValue.FindMember("k1").The third optimization is the at_pointer interface simdjson provided, this could directly get the json field from original document.	2022-08-16 14:49:50 +08:00
Ashin Gau	0b9bfd15b7	[feature-wip](parquet-reader) parquet physical type to doris logical type (#11769 ) Two improvements have been added: 1. Translate parquet physical type into doris logical type. 2. Decode parquet column chunk into doris ColumnPtr, and add unit tests to show how to use related API.	2022-08-15 16:08:11 +08:00
Ashin Gau	8f5aed27ec	[feature-wip](parquet-reader)read and decode parquet physical type (#11637 ) # Proposed changes Read and decode parquet physical type. 1. The encoding type of boolean is bit-packing, this PR introduces the implementation of bit-packing from Impala 2. Create a parquet including all the primitive types supported by hive ## Remaining Problems 1. At present, only physical types are decoded, and there is no corresponding and conversion methods with doris logical. 2. No parsing and processing Decimal type / Timestamp / Date. 3. Int_8 / Int_16 is stored as Int_32. How to resolve these types.	2022-08-11 10:17:32 +08:00
Kang	f9b151744d	optimize topn query if order by columns is prefix of sort keys of table (#10694 ) * [feature](planner): push limit to olapscan when meet sort. * if olap_scan_node's sort_info is set, push sort_limit, read_orderby_key and read_orderby_key_reverse for olap scanner * There is a common query pattern to find latest time serials data. eg. SELECT * from t_log WHERE t>t1 AND t<t2 ORDER BY t DESC LIMIT 100 If the ORDER BY columns is the prefix of the sort key of table, it can be greatly optimized to read much fewer data instead of read all data between t1 and t2. By leveraging the same order of ORDER BY columns and sort key of table, just read the LIMIT N rows for each related segment and merge N rows. 1. set read_orderby_key to true for read_params and _reader_context if olap_scan_node's sort info is set. 2. set read_orderby_key_reverse to true for read_params and _reader_context if is_asc_order is false. 3. rowset reader force merge read segments if read_orderby_key is true. 4. block reader and tablet reader force merge read rowsets if read_orderby_key is true. 5. for ORDER BY DESC, read and compare in reverse order 5.1 segment iterator read backward using a new BackwardBitmapRangeIterator and reverse the result block before return to caller. 5.2 VCollectIterator::LevelIteratorComparator, VMergeIteratorContext return opposite result for _is_reverse order in its compare function. Co-authored-by: jackwener <jakevingoo@gmail.com>	2022-08-09 09:08:44 +08:00
Ashin Gau	37d1180cca	[feature-wip](parquet-reader)decode parquet data (#11536 )	2022-08-08 12:44:06 +08:00
slothever	e8a344b683	[feature-wip](parquet-reader) add predicate filter and column reader (#11488 )	2022-08-08 10:21:24 +08:00
Ashin Gau	44a1a20e65	[feature-wip](parquet-reader)parse parquet schema (#11381 ) Analyze schema elements in parquet FileMetaData, and generate the hierarchy of nested fields. For exmpale: 1. primitive type ``` // thrift: optional int32 <column-name>; // sql definition: <column-name> int32; ``` 2. nested type ``` // thrift: optional group <column-name> (LIST) { repeated group bag { optional group array_element (LIST) { repeated group bag { optional int32 array_element } } } } // sql definition: <column-name> array<array<int32>> ```	2022-08-02 10:56:13 +08:00
slothever	e4bc3f6b6f	[feature-wip] (parquet-reader) add parquet reader impl template (#11285 )	2022-07-29 14:30:31 +08:00
Gabriel	328a225050	[feature-wip] (datetimev2) support window funnel and modify valid dat… (#11277 ) * [feature-wip] (datetimev2) support window funnel and modify valid date range	2022-07-28 14:06:26 +08:00
Pxl	1b4a2c287e	[Improvement][chore] replace from_decv2_to_packed128 to decv2.value (#11261 )	2022-07-28 10:41:27 +08:00
Gabriel	72d2feae99	[feature-wip] Support all date functions for datev2/datetimev2 (#11265 ) * [feature-wip] (datetimev2) support convert_tz function * [feature-wip] Support all date functions for datev2/datetimev2	2022-07-28 08:18:59 +08:00
Gabriel	d67029c830	[feature-wip] (datetimev2) support `cast` between datetimev2 with different scales (#11198 ) * [feature-wip] (datetimev2) support `cast` between datetimev2 with different scale	2022-07-26 22:36:13 +08:00
Gabriel	823088a9eb	[FOLLOW-UP] (datetimev2) complete date function ut and built-in function declaration (#11154 )	2022-07-26 17:48:57 +08:00
Gabriel	829d534e12	[Improvement] Replace `switch` with `constexpr` to boost date functions (#11134 )	2022-07-23 22:58:59 +08:00
Gabriel	babab5d535	[feature-wip] support datetimev2 (#11085 )	2022-07-23 16:07:59 +08:00
Xinyi Zou	4960043f5e	[enhancement] Refactor to improve the usability of MemTracker (step2) (#10823 )	2022-07-21 17:11:28 +08:00
yixiutt	dc2b709f6f	[Bug](compaction) fix uniq key compaction bug that does not count merged rows right(#10971 ) When a rowset includes multiple segments, segments rows will be merged in generic_iterator but merged_rows is not maintained. Compaction will failed in check_correctness. Co-authored-by: yixiutt <yixiu@selectdb.com>	2022-07-20 12:07:45 +08:00
camby	09d19e3f0f	[feature-wip](array-type) explode support more sub types (#10673 ) 1. explode support more sub types; 2. explode support nullable elements; Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-07-17 18:08:30 +08:00
zxealous	5793cb11d0	[feature-wip] (array-type) function concat_ws support array (#10749 ) Issue #10052 function concat_ws support array	2022-07-17 17:50:39 +08:00
camby	00c9455f16	[fix](array-type) fix arrow column to doris array column (#10855 ) * support merge array column, while convert from arrow column to doris array column * fix typo Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-07-16 11:49:42 +08:00
Gabriel	dc6fbcce14	[feature-wip] (datev2) modify datev2 format in memory (#10873 ) * [feature-wip] (datev2) modify datev2 format in memory * update	2022-07-15 19:57:38 +08:00
carlvinhust2012	1112dba525	[be ut]add some case for array type in block_test (#10656 ) Co-authored-by: hucheng01 <hucheng01@baidu.com>	2022-07-09 12:00:42 +08:00
camby	fe8acdb268	[feature-wip](array-type) add agg function collect_list and collect_set (#10606 ) add codes for collect_list and collect_set and update regression output, before output format for ARRAY(string) already changed. Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-07-08 12:48:46 +08:00
Dongyang Li	8012d63ea0	[fix] substr('', 1, 5) return empty string instead of null (#10622 )	2022-07-06 22:51:02 +08:00
camby	1f57fcc4e9	remove duplicate codes from function_test_util.cpp (#10607 ) Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-07-05 20:43:56 +08:00
camby	ec6620ae3e	[feature-wip](array-type) add function arrays_overlap (#10233 )	2022-06-30 08:12:29 +08:00
Gabriel	ca94867b4e	[Feature-wip] add date v2 type (#9916 )	2022-06-26 16:07:56 +08:00
Adonis Ling	5e47b03595	[feature-wip](array-type) Add array aggregation functions (#10108 )	2022-06-17 11:07:49 +08:00
Xinyi Zou	d58e00c49c	[fix](brpc) Embed serialized request into the attachment and transmit it through http brpc (#9803 ) When the length of `Tuple/Block data` is greater than 2G, serialize the protoBuf request and embed the `Tuple/Block data` into the controller attachment and transmit it through http brpc. This is to avoid errors when the length of the protoBuf request exceeds 2G: `Bad request, error_text=[E1003]Fail to compress request`. In #7164, `Tuple/Block data` was put into attachment and sent via default `baidu_std brpc`, but when the attachment exceeds 2G, it will be truncated. There is no 2G limit for sending via `http brpc`. Also, in #7921, consider putting `Tuple/Block data` into attachment transport by default, as this theoretically reduces one serialization and improves performance. However, the test found that the performance did not improve, but the memory peak increased due to the addition of a memory copy.	2022-06-13 20:41:48 +08:00
camby	6fab1cbf3c	[feature-wip](array-type) Add array functions size and cardinality (#9921 ) Co-authored-by: cambyzju <zhuxiaoli01@baidu.com>	2022-06-09 15:03:03 +08:00
yinzhijian	19bc14cf8d	[feature-wip](array-type) Add array type support for vectorized parquet-orc scanner (#9856 ) Only support one level array now. for example: - nullable(array(nullable(tinyint))) is support. - nullable(array(nullable(array(xx))) is not support.	2022-06-09 12:11:47 +08:00
HappenLee	94089b9192	[Refactor] Use file factory to replace create file reader/writer (#9505 ) 1. Simplify code logic and improve abstraction 2. Fix the mem leak of raw pointer Co-authored-by: lihaopeng <lihaopeng@baidu.com>	2022-06-08 15:07:39 +08:00
Pxl	c0ad1be1bd	[Enhancement][Chore] remove breakpad and unused variable (#9937 )	2022-06-02 20:52:17 +08:00
Gabriel	632f7a3d3d	[Feature] add `weekday` function on vectorized engine (#9901 )	2022-06-01 14:47:37 +08:00
HappenLee	0cba6b7d95	[Bug][Fix] One Rowset have same key output in unique table (#9858 ) Co-authored-by: lihaopeng <lihaopeng@baidu.com>	2022-05-31 12:29:16 +08:00
Adonis Ling	f377c26bf7	[refactor][be] Optimize headers (#9708 )	2022-05-30 16:12:10 +08:00
Jing Shen	7b98dd438d	[feature](function) Add nvl function (#9726 )	2022-05-30 09:43:00 +08:00
yinzhijian	cbbda7857b	[feature-wip](parquet-orc) Support orc scanner in vectorized engine (#9541 )	2022-05-26 21:39:12 +08:00
Pxl	13c1d20426	[Bug] [Vectorized] add padding when load char type data (#9734 )	2022-05-26 16:51:01 +08:00
Gabriel	8470543144	[Improvement] fix typo (#9743 )	2022-05-25 19:29:01 +08:00
xiepengcheng01	31e40191a8	[Refactor] add vpre_filter_expr for vectorized to improve performance (#9508 )	2022-05-22 11:45:57 +08:00
HappenLee	8fa677b59c	[Refactor][Bug-Fix][Load Vec] Refactor code of basescanner and vjson/vparquet/vbroker scanner (#9666 ) * [Refactor][Bug-Fix][Load Vec] Refactor code of basescanner and vjson/vparquet/vbroker scanner 1. fix bug of vjson scanner not support `range_from_file_path` 2. fix bug of vjson/vbrocker scanner core dump by src/dest slot nullable is different 3. fix bug of vparquest filter_block reference of column in not 1 4. refactor code to simple all the code It only changed vectorized load, not original row based load. Co-authored-by: lihaopeng <lihaopeng@baidu.com>	2022-05-20 11:43:03 +08:00
yinzhijian	bee5c2f8aa	[feature-wip](parquet-vec) Support parquet scanner in vectorized engine (#9433 )	2022-05-17 09:37:17 +08:00
zhangstar333	953429e370	[fix](function) fix last_value get wrong result when have order by clause (#9247 )	2022-05-15 23:56:01 +08:00
carlvinhust2012	b817efd652	[feature] add vectorized vjson_scanner (#9311 ) This pr is used to add the vectorized vjson_scanner, which can support vectorized json import in stream load flow.	2022-05-14 09:50:05 +08:00

1 2

94 Commits