doris

Author	SHA1	Message	Date
starocean999	c4341d3d43	[fix](like)prevent null pointer by unimplemented like_vec functions (#12910 ) * [fix](like)prevent null pointer by unimplemented like_vec functions * fix pushed like predicate on dict encoded column bug	2022-09-27 10:02:10 +08:00
pengxiangyu	e040dccbec	[fix](remote)fix bug for delete s3 dir and list s3 dir (#12918 ) * fix bug for delete s3 dir and list s3 dir	2022-09-27 09:54:37 +08:00
Xinyi Zou	b14b178928	[enhancement](memory) Trigger load channel flush based on process physical memory to avoid OOM #12960 When the physical memory of the process reaches 90% of the mem limit, trigger the load channel mgr to brush down The default value of be.conf mem_limit is changed from 90% to 80%, and stability is the priority. Fix deadlock in arena_locks in BufferPool::BufferAllocator::ScavengeBuffers and _lock in DebugString	2022-09-27 09:07:38 +08:00
Pxl	12d6efa92b	[Bug](function) fix substr return null on row-based engine #12906	2022-09-27 08:47:32 +08:00
Xiaocc	5790d23624	[fix](transfer_thread) fix the loss of notification. (#12988 )	2022-09-27 08:44:02 +08:00
Pxl	8731eea26e	[Chore](clang) fix some build fail on clang15 (#12882 ) remove unused variables	2022-09-26 23:13:28 +08:00
zxealous	595a5337dc	fix doc typos (#12967 )	2022-09-26 20:11:26 +08:00
Shane	35076431ab	[fix](column)fix get_shrinked_column misspell (#12961 ) Fix misspell	2022-09-26 17:32:03 +08:00
TengJianPing	1bb42a7bc0	[function](hash) add support of murmur_hash3_64 (#12923 )	2022-09-26 14:23:37 +08:00
Xinyi Zou	72220440dc	[fix](memtracker) Remove mem tracker record mem pool actual memory usage #12954 In order to avoid different mem tracker consumption values of multiple queries/loads, and the difference between the virtual memory of alloc and the physical memory actually increased by the process. The memory alloc in PODArray and mempool will not be recorded in the query/load mem tracker immediately, but will be gradually recorded in the mem tracker during the memory usage. But mem pool allocates memory from chunk allocator. If this chunk is used after the second time, it may have used physical memory. The above mechanism will cause the load channel memory statistics to be less than the actual value.	2022-09-26 12:54:06 +08:00
luozenglin	0fcb93aae2	[fix](parquet) fix write error data as parquet format. (#12864 ) * [fix](parquet) fix write error data as parquet format. Fix incorrect data conversion when writing tiny int and small int data to parquet files in non-vectorized engine.	2022-09-26 10:41:17 +08:00
Tiewei Fang	acd5d67355	[feature-wip](new-scan)Add new odbc scanner and new odbc scan node (#12899 )	2022-09-26 09:24:25 +08:00
Jerry Hu	56fc00cb53	[chore](config) increase minimum thread num of some thread pool (#12917 ) Too small minimum thread num will cause additional overhead for creating and recycling threads.	2022-09-26 09:00:18 +08:00
Adonis Ling	32144ccda8	[Enhancement](debugging) Add more debug info for clang build (#12845 )	2022-09-26 08:50:12 +08:00
Ashin Gau	692176ec07	[feature-wip](parquet-reader) pre read page data in advance to avoid frequent seek (#12898 ) 1. Fix the bug of file position in `HdfsFileReader` 2. Reserve enough buffer for `ColumnColumnReader` to read large continuous memory	2022-09-25 21:21:06 +08:00
Gabriel	380c3f42ab	[Refactor](datev2) Update comments for datev2/datetimev2 (#12823 )	2022-09-25 18:43:32 +08:00
Gabriel	f879a51ce9	[Improvement](dict) optimize dictionary column (#12852 )	2022-09-25 18:29:10 +08:00
Gabriel	d8e8bc0e69	[Improvement](predicate) Replace for-loop by memcpy (#12867 )	2022-09-25 18:27:59 +08:00
Shane	59699a4321	[feature](JSON datatype)Support JSON datatype (#10322 ) Add `JSON` datatype, following features are implemented by this PR: 1. `CREATE` tables with `JSON` type columns 2. `INSERT` values containing `JSON` type value stored in `String`, which is represented as binary format(AKA `JSONB`) at BE 3. `SELECT` JSON columns Detail design refers [DSIP-016: Support JSON type](https://cwiki.apache.org/confluence/display/DORIS/DSIP-016%3A+Support+JSON+type) * add JSONB data storage format type * fix JsonLiteral resolve bug * add DataTypeJson case in data_type_factory * add JSON syntax check in FE * add operators for jsonb_document, currently not support comparison between any JSON type value * add ColumnJson and DataTypeJson * add JsonField to store JsonValue * add JsonValue to convert String JSON to BINARY JSON and JsonLiteral case for vliteral * add push_json for MysqlResultWriter * JSON column need no zone_map_index * Revert "JSON column need no zone_map_index" This reverts commit f71d1ce1ded9dbae44a5d58abcec338816b70d79. * add JSON writer and reader, ignore zone-map for JSON column * add json_to_string for DataTypeJson * add olap_data_convertor for JSON type * add some enum * add OLAP_FIELD_TYPE_JSON type, FieldTypeTraits for it and corresponding cases or functions * fix column_json offsets overflow bug, format code * remove useless TODOs, add CmpType cases for JSON type * add license header * format license * format be codes * resolve rebase master conflicts * fix bugs for CREATE and meta related code * refactor JsonValue constructors, add fe JSON cases and fix some bugs, reformat codes * modification be codes along code review advice * fix rebase conflicts with master * add unit test for json_value and column_json * fix rebase error * rename json to jsonb * fix some data convert bugs, set Mysql type to JSON	2022-09-25 14:06:49 +08:00
zhannngchen	57d5f69814	[fix](load) print detailed error message (#12938 ) fix flush failure return message	2022-09-25 10:31:41 +08:00
starocean999	dd6ed5a9a7	[fix](function)fix string split function buffer overflow (#12834 )	2022-09-24 17:32:00 +08:00
Jibing-Li	f1a64ea09f	[fix](new-scan)Fix new scanner load job bugs (#12903 ) Fix bugs: 1. Fe need to send file format (e.g. parquet, orc ...) to be while processing load jobs using new scanner. 2. Try to get parquet file column type from SchemaElement.type before getting from Logical type and Converted type.	2022-09-24 17:21:19 +08:00
zhannngchen	3bb920ba54	[Enhancement](load) Refine the load channel flush policy on mem limit (#12716 ) 1. Remove single load channel mem limit, only use load channel mgr mem limit 2. Default load channel mgr mem limit from 50% to 80% 3. load channel mgr add soft mem limit. When the soft limit is exceeded, other threads will not hang, only current thread triggers flush 4. When exceed load channel mgr mem limit, find a load channel with the largest mem usage, continue to find a tablet channel with the largest mem usage, and try to flush 1/3 of the mem usage of this tablet channel.	2022-09-24 10:01:13 +08:00
yiguolei	7b230e41a8	[bugfix](scanner) olap scanner compute is wrong (#12857 ) Co-authored-by: yiguolei <yiguolei@gmail.com>	2022-09-24 09:59:59 +08:00
Xinyi Zou	34d6d36ff5	fix transfer to tracker (#12932 ) ~MemTrackerLimiter() repeated consumption of _untracked_mem, resulting in inaccurate process mem tracker.	2022-09-24 09:01:05 +08:00
Yongqiang YANG	9dc35ab534	[fix](streamload) set coord for streamLoad (#12744 ) When a stream load is canceled, status is reported to coord.	2022-09-23 20:23:19 +08:00
Ashin Gau	5bfdfac387	[feature-wip](parquet-reader) add parquet reader profile (#12797 ) Add profile for parquet reader. New counters: - ParquetFilteredGroups: Filtered row groups by `RowGroup` min-max statistics - ParquetReadGroups: The number of row groups to read - ParquetFilteredRowsByGroup: The number of filtered rows by `RowGroup` min-max statistics - ParquetFilteredRowsByPage: The number of filtered rows by page min-max statistics - ParquetFilteredBytes: The filtered bytes by `RowGroup` min-max statistics - ParquetReadBytes: The total bytes in `ParquetReadGroups`, may be further filtered If a page is skipped as a whole ## Result ``` ┌──────────────────────────────────────────────────────┐ │[0: VFILE_SCAN_NODE] │ │(Active: 1s29ms, non-child: 96.42) │ │ - Counters: │ │ - BytesRead: 0.00 │ │ - FileReadCalls: 1.826K (1826) │ │ - FileReadTime: 510.627ms │ │ - FileRemoteReadBytes: 65.23 MB │ │ - FileRemoteReadCalls: 1.146K (1146) │ │ - FileRemoteReadRate: 128.29331970214844 MB/sec │ │ - FileRemoteReadTime: 508.469ms │ │ - NumDiskAccess: 0 │ │ - NumScanners: 1 │ │ - ParquetFilteredBytes: 0.00 │ │ - ParquetFilteredGroups: 0 │ │ - ParquetFilteredRowsByGroup: 0 │ │ - ParquetFilteredRowsByPage: 6.600003M (6600003)│ │ - ParquetReadBytes: 2.13 GB │ │ - ParquetReadGroups: 20 │ │ - PeakMemoryUsage: 0.00 │ │ - PredicateFilteredRows: 3.399797M (3399797) │ │ - PredicateFilteredTime: 133.302ms │ │ - RowsRead: 3.399997M (3399997) │ │ - RowsReturned: 200 │ │ - RowsReturnedRate: 194 │ │ - TotalRawReadTime(*): 726.566ms │ │ - TotalReadThroughput: 0.0 /sec │ │ - WaitScannerTime: 1s27ms │ └──────────────────────────────────────────────────────┘ ```	2022-09-23 18:42:14 +08:00
HappenLee	f7e3ca29b5	[Opt](Vectorized) Support push down no grouping agg (#12803 ) Support push down no grouping agg	2022-09-23 18:29:54 +08:00
Yongqiang YANG	a7d42b5d81	[fix](streamload&sink) release and allocate memory in the same tracker (#12820 ) 1. HttpServer threads allocate bytebuffer and put them into streamload pipe, but scanner thread release them with query tracker. 2. We can assume brpc allocate memory in doris thread. Above problems leads to wrong result of memtracker.	2022-09-23 17:51:44 +08:00
zhangstar333	617820b1f5	[Refactor](parquet) refactor parquet write to uniform and consistent logic (#12730 )	2022-09-23 09:12:34 +08:00
Zhengguo Yang	8fcd8ed8b3	[chore](build) add option to disable -frecord-gcc-switches (#12846 )	2022-09-22 15:38:14 +08:00
xueweizhang	70ab9cb43e	[feature](http) refactor version info and add new http api for get version info (#12513 ) Refactor version info and add new http api for get version info	2022-09-22 10:53:04 +08:00
Yongqiang YANG	77e423042c	(brpc) donot use pooled brpc (#12754 ) It seems that pooled brpc does not release port timely.	2022-09-22 10:00:26 +08:00
starocean999	57b3c03371	[enhancement](like)pass data to like function in block not in row (#12825 ) The like predicate process data in block perform better than in row. Currently, only not null column is optimized, nullable column will be handled later. SELECT COUNT(*) FROM hits WHERE URL LIKE '%google%'; before: ~680ms after: ~570ms	2022-09-22 09:59:30 +08:00
yiguolei	32551a7263	[bugfix](predicate column) data maybe wrong if not a single page (#12796 ) Co-authored-by: yiguolei <yiguolei@gmail.com>	2022-09-22 09:55:31 +08:00
Jibing-Li	4b95b4e41d	[feature-wip](file-scanner)Get column type from parquet schema (#12833 ) Get schema from parquet reader. The new VFileScanner need to get file schema (column name to type map) from parquet file while processing load job, this pr is to set the type information for parquet columns.	2022-09-22 09:35:37 +08:00
slothever	1ca6d559e4	[feature-wip](parquet-reader) refactor some arguments for parquet reader (#12771 ) refactor some arguments for parquet reader 1. Add new parquet context to wrap reader arguments 2. Reduced some arguments for function call Co-authored-by: jinzhe <jinzhe@selectdb.com>	2022-09-22 09:34:01 +08:00
Gabriel	e21ffac419	[Improvement](dateformat) Improve efficiency for function `date_format` (#12811 )	2022-09-21 22:38:16 +08:00
Jibing-Li	fbdebe2424	[feature-wip](new-scan)Add load counter for VFileScanner (#12812 ) The new scanner (VFileScanner) need a counter to record two values in load job. 1. The number of rows unselected by pre-filter, and 2. The number of rows filtered by unmatched schema or other error. This pr is to implement the counter.	2022-09-21 20:59:13 +08:00
Xinyi Zou	c55d08fa2f	[fix](memtracker) Refactor load channel mem tracker to improve accuracy (#12791 ) The mem hook record tracker cannot guarantee that the final consumption is 0, nor can it guarantee that the memory alloc and free are recorded in a one-to-one correspondence. In the life cycle of a memtable from insert to flush, the memory free of hook is more than that of alloc, resulting in tracker consumption less than 0. In order to avoid the cumulative error of the upper load channel tracker, the memtable tracker consumption is reset to zero on destructor.	2022-09-21 20:16:19 +08:00
Xinyi Zou	b41eaa5ac0	[fix](memtracker) Introduce orphan mem tracker to verify memory tracking accuracy (#12794 ) The mem hook consumes the orphan tracker by default. If the thread does not attach other trackers, by default all consumption will be passed to the process tracker through the orphan tracker. In real time, consumption of all other trackers + orphan tracker consumption = process tracker consumption. Ideally, all threads are expected to attach to the specified tracker, so that "all memory has its own ownership", and the consumption of the orphan mem tracker is close to 0, but greater than 0.	2022-09-21 15:47:10 +08:00
Jerry Hu	8f4bb0f804	[improvement](agg) iterate aggregation data in memory written order (#12704 ) Following the iteration order of the hash table will result in out-of-order access to aggregate states, which is very inefficient. Traversing aggregate states in memory write order can significantly improve memory read efficiency. Test hash table items count: 3.35M Before this optimization: insert keys into column takes 500ms With this optimization only takes 80ms	2022-09-21 14:58:50 +08:00
zhannngchen	27f7ae258d	[Enhancement](load) optimize flush policy to avoid small segments #12706 In current policy, if mem-limit exceeded, load channel will pick tablets that consume most memory, but mem_consumption contains memory in flush, if some delta writer flushing a full memtable(default 200MB), the current memtable might be very small, we should avoid flush such memtable, which can generate a very small segment.	2022-09-21 14:33:05 +08:00
Jibing-Li	ec2b3bf220	[feature-wip](new-scan)Refactor VFileScanner, support broker load, remove unused functions in VScanner base class. (#12793 ) Refactor of scanners. Support broker load. This pr is part of the refactor scanner tasks. It provide support for borker load using new VFileScanner. Work still in progress.	2022-09-21 12:49:56 +08:00
luozenglin	b6e20db997	[fix](outfile) select OBJECT and HLL columns into outfile as null. (#12734 )	2022-09-21 11:24:31 +08:00
Gabriel	3cfaae0031	[Improvement](sort) Use heap sort to optimize sort node (#12700 )	2022-09-21 10:01:52 +08:00
Xin Liao	a5643822de	[feature-wip](unique-key-merge-on-write) fix calculate delete bitmap when has sequence column (#12789 ) when the rowset has multiple segments with sequence column, we should compare sequence id with previous segment.	2022-09-21 09:21:07 +08:00
Xinyi Zou	bd4bfa8f00	[fix](memtracker) Fix thread mem tracker try consume accuracy #12782	2022-09-21 09:20:41 +08:00
AlexYue	c72a19f410	[BugFix](VExprContext) capture error status to prevent incorrect func call which causes coredump #12779	2022-09-21 09:20:16 +08:00
AlexYue	f1539761e8	[Bugfix](string_functions) rearrange code to avoid global buffer overflow in FindInSetOp::execute (#12677 )	2022-09-21 09:19:38 +08:00

1 2 3 4 5 ...

2855 Commits