Skip to content

[fix](score) disable search topn with extra predicates - #65821

Merged
airborne12 merged 2 commits into
apache:masterfrom
LIANG751234313:fix/search-score-topn-extra-predicates
Aug 21, 2026
Merged

[fix](score) disable search topn with extra predicates#65821
airborne12 merged 2 commits into
apache:masterfrom
LIANG751234313:fix/search-score-topn-extra-predicates

Conversation

@LIANG751234313

@LIANG751234313 LIANG751234313 commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

What problem does this PR solve?

Issue Number: close #xxx

Related PR: #xxx

Problem Summary:
Search score TopN pushdown may return incorrect results when the search predicate is combined with additional predicates, because the pushed TopN limit can be applied before the remaining predicates are evaluated. This change disables the pushed search TopN limit in those cases while preserving the virtual score column pushdown, and adds regression coverage for search predicates combined with equality, range, match, score range and multiple search predicates.

Release note

None

Check List (For Author)

  • Test

    • Regression test
    • Unit Test
    • Manual test (add detailed scripts or steps below)
    • No need to test or manual test. Explain why:
      • This is a refactor/code format and no logic has been changed.
      • Previous test can cover this change.
      • No code files have been changed.
      • Other reason
  • Behavior changed:

    • No.
    • Yes.
  • Does this need documentation?

    • No.
    • Yes.

Check List (For Reviewer who merge this PR)

  • Confirm the release note
  • Confirm test cases
  • Confirm document
  • Add branch pick label

@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@morrySnow morrySnow changed the title [fix](fe) disable search topn with extra predicates [fix](score) disable search topn with extra predicates Jul 20, 2026
airborne12
airborne12 previously approved these changes Jul 30, 2026

@airborne12 airborne12 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@airborne12

Copy link
Copy Markdown
Member

run buildall

@github-actions github-actions Bot added the approved Indicates a PR has been approved by one committer. label Jul 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

PR approved by at least one committer and no changes requested.

@github-actions

Copy link
Copy Markdown
Contributor

PR approved by anyone and no changes requested.

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 29747 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit 94d0c144a0ac48ced6cf672a568bb239c596454e, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17675	4168	4115	4115
q2	2036	319	205	205
q3	10276	1431	828	828
q4	4682	475	346	346
q5	7497	879	573	573
q6	197	183	149	149
q7	782	819	622	622
q8	9340	1575	1510	1510
q9	5538	4326	4331	4326
q10	6735	1779	1468	1468
q11	518	365	340	340
q12	739	605	471	471
q13	18123	3542	2799	2799
q14	266	267	244	244
q15	q16	789	776	708	708
q17	1005	1034	1030	1030
q18	6846	5636	5542	5542
q19	1150	1232	1082	1082
q20	831	695	625	625
q21	5804	2591	2452	2452
q22	438	365	312	312
Total cold run time: 101267 ms
Total hot run time: 29747 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	4537	4433	4495	4433
q2	301	320	220	220
q3	4570	4954	4383	4383
q4	2100	2166	1384	1384
q5	4459	4343	4359	4343
q6	240	187	135	135
q7	1755	2003	1859	1859
q8	2590	2247	2191	2191
q9	8099	8137	7815	7815
q10	4764	4666	4219	4219
q11	576	423	419	419
q12	773	779	562	562
q13	3216	3642	2983	2983
q14	297	307	284	284
q15	q16	715	727	637	637
q17	1349	1342	1309	1309
q18	8072	7426	7471	7426
q19	1202	1152	1103	1103
q20	2236	2205	1931	1931
q21	5397	4602	4457	4457
q22	516	466	430	430
Total cold run time: 57764 ms
Total hot run time: 52523 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 179221 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit 94d0c144a0ac48ced6cf672a568bb239c596454e, data reload: false

query5	4318	641	492	492
query6	467	235	208	208
query7	4887	611	359	359
query8	350	190	183	183
query9	8771	4099	4101	4099
query10	512	360	301	301
query11	5850	2375	2146	2146
query12	160	106	101	101
query13	1262	609	446	446
query14	6258	5244	4950	4950
query14_1	4233	4236	4231	4231
query15	220	204	177	177
query16	1067	478	460	460
query17	931	698	565	565
query18	2431	488	338	338
query19	216	186	146	146
query20	112	106	108	106
query21	233	158	134	134
query22	13563	13522	13364	13364
query23	17549	16481	16161	16161
query23_1	16313	16342	16227	16227
query24	7400	1797	1296	1296
query24_1	1336	1289	1277	1277
query25	540	431	378	378
query26	1339	361	215	215
query27	2603	663	388	388
query28	4505	1992	2012	1992
query29	1066	602	481	481
query30	356	278	229	229
query31	1130	1113	988	988
query32	100	61	61	61
query33	520	356	262	262
query34	1207	1159	670	670
query35	785	788	673	673
query36	1030	1045	871	871
query37	155	109	94	94
query38	1907	1734	1672	1672
query39	880	903	875	875
query39_1	834	848	869	848
query40	244	158	141	141
query41	65	62	64	62
query42	104	94	92	92
query43	322	327	292	292
query44	1453	795	770	770
query45	192	177	171	171
query46	1053	1249	762	762
query47	2193	2157	2043	2043
query48	405	413	309	309
query49	578	428	307	307
query50	1149	434	338	338
query51	10577	10583	10475	10475
query52	91	88	77	77
query53	266	284	220	220
query54	277	237	220	220
query55	74	72	66	66
query56	297	300	312	300
query57	1310	1287	1213	1213
query58	289	267	253	253
query59	1577	1651	1463	1463
query60	320	275	255	255
query61	150	152	151	151
query62	539	498	431	431
query63	242	199	202	199
query64	2896	1229	1004	1004
query65	4870	4726	4752	4726
query66	1875	517	401	401
query67	30073	29776	29623	29623
query68	3121	1522	1033	1033
query69	434	314	274	274
query70	937	822	833	822
query71	393	350	342	342
query72	3291	2891	2301	2301
query73	854	782	425	425
query74	5094	5006	4730	4730
query75	2551	2551	2147	2147
query76	2319	1197	764	764
query77	343	381	268	268
query78	12311	12241	11621	11621
query79	1393	1189	710	710
query80	1315	556	488	488
query81	524	342	290	290
query82	623	156	119	119
query83	404	330	307	307
query84	326	157	131	131
query85	986	667	532	532
query86	413	246	234	234
query87	1861	1840	1781	1781
query88	3833	2842	2844	2842
query89	431	369	334	334
query90	2043	202	190	190
query91	206	194	165	165
query92	59	57	54	54
query93	1607	1610	982	982
query94	741	357	311	311
query95	809	596	491	491
query96	1138	792	363	363
query97	2629	2624	2474	2474
query98	213	215	203	203
query99	1074	1117	989	989
Total cold run time: 265204 ms
Total hot run time: 179221 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 24.85 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit 94d0c144a0ac48ced6cf672a568bb239c596454e, data reload: false

query1	0.01	0.01	0.01
query2	0.09	0.05	0.04
query3	0.25	0.13	0.14
query4	1.60	0.14	0.14
query5	0.24	0.22	0.23
query6	1.16	0.84	0.82
query7	0.04	0.01	0.01
query8	0.06	0.04	0.04
query9	0.41	0.32	0.32
query10	0.56	0.55	0.56
query11	0.19	0.14	0.14
query12	0.17	0.14	0.14
query13	0.48	0.48	0.48
query14	1.03	1.02	1.01
query15	0.64	0.59	0.60
query16	0.32	0.33	0.33
query17	1.08	1.12	1.07
query18	0.23	0.22	0.21
query19	2.06	2.01	1.95
query20	0.02	0.01	0.02
query21	15.45	0.21	0.15
query22	4.94	0.06	0.05
query23	16.12	0.31	0.12
query24	2.93	0.43	0.33
query25	0.10	0.06	0.03
query26	0.74	0.21	0.15
query27	0.03	0.04	0.03
query28	3.52	0.91	0.56
query29	12.53	4.13	3.27
query30	0.27	0.15	0.15
query31	2.77	0.60	0.31
query32	3.24	0.59	0.49
query33	3.23	3.17	3.20
query34	15.40	4.22	3.54
query35	3.56	3.52	3.50
query36	0.55	0.44	0.42
query37	0.09	0.06	0.06
query38	0.05	0.04	0.03
query39	0.04	0.03	0.03
query40	0.17	0.16	0.16
query41	0.08	0.03	0.03
query42	0.04	0.03	0.03
query43	0.04	0.04	0.03
Total cold run time: 96.53 s
Total hot run time: 24.85 s

@hello-stephen

Copy link
Copy Markdown
Contributor

FE Regression Coverage Report

Increment line coverage 84.62% (11/13) 🎉
Increment coverage report
Complete coverage report

@airborne12

Copy link
Copy Markdown
Member

/review

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

Request changes. The patch correctly disables early SEARCH Top-K when the logical filter contains explicit residual predicates, and the zero sentinel remains present through FE translation and switches the current/base BE paths to full SEARCH and score collection. However, the preserved lone-SEARCH branch is still unsafe on supported unique-key MOW reads: SEARCH selects K physical documents before the segment delete bitmap removes hidden versions, so a deleted highest scorer can suppress the next visible row. That is a reachable wrong-result path and needs to be addressed before merge.

The new regression suite is broad and its expected rankings/results are deterministic, but it does not cover the MOW visibility case. Two additional test gaps are called out inline conditionally: they matter only if the fix retains a selective positive-limit/recursive-classifier branch.

Review checkpoints

  • Goal and correctness: Explicit extra-predicate cases are substantially fixed, but the supported storage-visibility path above leaves the goal incomplete.
  • Scope and maintainability: The production change is small and focused; the helper is readable and reuses existing expression traversal.
  • Concurrency and lifecycle: No new locking, shared mutable state, thread-lifecycle, or static-initialization behavior is introduced.
  • Configuration and compatibility: No configuration, protocol, or storage-format change is introduced. The existing optional score-limit field and zero semantics are compatible with the target-base BE path checked for this review.
  • Parallel execution paths: Direct/nested SEARCH, MATCH-only scoring, score-range extraction, LIMIT/OFFSET, multiple SEARCH predicates, row ranges, delete conditions, and MOW delete bitmaps were traced. MATCH-only and the explicit predicate classifier are sound; post-SEARCH visibility is the blocking gap.
  • Transactions, persistence, and writes: No write/transaction path changes, but read visibility for committed MOW versions is affected by the blocking issue.
  • Error handling and observability: No new exception/nullability issue was found. Existing EXPLAIN output exposes SCORE SORT LIMIT, which can test the material positive-versus-zero planner choice if that choice remains.
  • Performance: Zero intentionally uses full document/score collection, so it costs more work but is the necessary correctness fallback until a positive limit can account for all later masks.
  • Tests and CI: Changed result files match the suite and exercise indexed/plain predicates, MATCH, ranges, multiple/nested/negated SEARCH, score filtering, and OFFSET. Compile, CheckStyle, FE UT, and the relevant P0 regression checks are green. No local build or test was run because this review environment designates the bundle as review-only.
  • User focus: No additional review focus was provided.

if (!hasSearchPredicate) {
return false;
}
return nonScoreConjuncts.size() > 1 || !(nonScoreConjuncts.get(0) instanceof SearchExpression);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Account for storage visibility before keeping SEARCH Top-K

A lone SEARCH is not sufficient to make early Top-K safe. For example:

TopN(score DESC, LIMIT 1)
  Project(id, score() AS score)
    Filter(search('title:apple'))
      Scan(unique-key MOW table)

This branch keeps score_sort_limit = 1. In BE, SegmentIterator::_lazy_init() evaluates SEARCH in _get_row_ranges_by_column_conditions() before subtracting _opts.delete_bitmap. If a deleted old version has the highest score and a live row is second, SEARCH returns only the deleted row; the later bitmap subtraction removes it, and the upper TopN cannot recover the live runner-up. SEARCH is supported on MOW tables, so this is reachable even with no extra SQL conjunct. Please either disable the early limit whenever post-SEARCH visibility/range filters may exist, or apply those masks before SEARCH selects Top-K, and add a deleted-highest-score regression.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked the MOW case with EXPLAIN, and the actual plan does not keep a positive score_sort_limit for a SQL-level lone SEARCH on a unique-key MOW table.
For example:

SELECT id, score() AS s
FROM test_search_score_mow_topk_visibility
WHERE search('title:alpha')
ORDER BY s DESC
LIMIT 5;

produces:

SCORE SORT LIMIT: 0
PREDICATES: (search('title:alpha') AND (DORIS_DELETE_SIGN = 0))

So although the SQL text contains only one SEARCH predicate, the scan predicate is not a lone SEARCH in a unique-key table. Doris injects the hidden delete-sign predicate, and this PR classifies it as SEARCH plus an extra predicate, so the early SEARCH Top-K limit is disabled.

I agree that if a unique-key/MOW scan kept a positive score_sort_limit, it would be unsafe because SEARCH Top-K is evaluated before post-SEARCH visibility filters such as delete bitmap / delete-sign filtering. But that is not the actual plan produced by this PR for the MOW case. With SCORE SORT LIMIT: 0, BE takes the full doc-set collection path instead of collect_multi_segment_top_k(), and the upper TopN applies the final limit after visibility filtering.

I also tried to reproduce the deleted-highest-score case on a unique-key MOW table, comparing the SQL-level lone SEARCH query with a control query that forces SCORE SORT LIMIT: 0, and both returned the same correct visible rows. Could you provide a concrete MOW SQL/EXPLAIN where a unique-key MOW lone SEARCH still produces a positive SCORE SORT LIMIT? Otherwise, this P1 seems to be based on a plan shape that this PR does not generate.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The MOW plan is protected by the injected delete-sign predicate, but that does not close the broader visibility concern. DUP_KEYS provides a concrete counterexample because those tables do not receive __DORIS_DELETE_SIGN__.

For example, insert two SEARCH-matching rows where id 1 scores above id 2, then execute DELETE FROM t WHERE id = 1. The later score query still has the reduced plan

TopN(score DESC, LIMIT 1)
  Project(id, score() AS score)
    Filter(search('title:alpha'))
      Scan(DUP_KEYS table)

so this helper keeps SCORE SORT LIMIT: 1. BE attaches the legacy delete predicate to older rowsets in BetaRowsetReader, but SEARCH consumes IndexQueryContext::query_limit first in function_search.cpp; exact delete_condition_predicates evaluation happens later in SegmentIterator::_evaluate_short_circuit_predicate(). The deleted high scorer therefore consumes Top-1 and is then removed, while the live runner-up was never collected.

Please account for storage delete predicates (and any other late semantic mask) before retaining positive SEARCH Top-K, and add a deleted-highest-score DUP_KEYS regression. I am replying here instead of opening another inline thread because this is concrete evidence for the same invariant already raised above.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This PR is scoped to the case where SEARCH is combined with other planner-visible predicates in the same FE plan, and the Top-K limit may be applied before those remaining predicates are evaluated. The DUP_KEYS + DELETE case is a separate storage-side visibility problem: the delete predicate is applied later than the SEARCH Top-K limit, so the FE rewrite in this PR cannot safely reason about it in the same way.

I’ve filed a separate JIRA to track that broader bug: http://jira.selectdb-in.cc/browse/DORIS-28025.

sql "set enable_segment_limit_pushdown = true"
sql "set enable_inverted_index_query_cache = false"

qt_single_search """

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Cover the positive branch if it remains

These result-only queries cannot distinguish SCORE SORT LIMIT: 1 from the disable sentinel 0, because the upper TopN produces the same rows either way; the complete changed suite passes if the helper always returns zero. After addressing the visibility issue above, if any selective positive-limit branch remains, please add deterministic EXPLAIN assertions for that demonstrably safe case and for an extra-predicate zero case. If the fix removes positive SEARCH Top-K entirely, this assertion is unnecessary.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The added regression cases are primarily correctness tests. They are constructed so that, for SEARCH plus residual predicates, keeping a positive SCORE SORT LIMIT would prune the high-score row before the residual predicate is applied and produce the wrong result. Therefore the expected results verify the correctness issue that this PR fixes.

ORDER BY id
"""

qt_nested_search_with_other_predicate """

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Make the nested cases depend on recursive SEARCH detection

Neither named case fails with a shallow top-level-only classifier. In the OR case, category MATCH 'special' independently restores id 7 after SEARCH Top-1 keeps id 8; in the NOT case, the separate top-level search('body:beta') already causes the limit to be disabled. If the selective recursive classifier remains after the correctness fix, please make the alternative OR branch nonmatching and add a negated-only witness (or a direct FE rule test) so losing .collect(SearchExpression) changes the expected result.

}

long scoreLimit = topN.getLimit() + topN.getOffset();
long pushedScoreLimit = shouldDisableSearchTopN(filter.getConjuncts(), extractedScorePredicate)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please check overflow and do not push down like #64633
add a test case for it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Before that, the overflow has been checked:
if (Utils.addOverflows(topN.getLimit(), topN.getOffset())) { return null; }
If overflow occurs, an error will occur:“score() function requires WHERE clause with MATCH function, ORDER BY and LIMIT for optimization”,instead of “limit + offset overflows the long range”.
Do I need to change return null; to throw new AnalysisException("limit + offset overflows the long range");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The key is: return null (skip optimization) if no wrong result, else throw exception

@LIANG751234313
LIANG751234313 force-pushed the fix/search-score-topn-extra-predicates branch from 94d0c14 to a96afba Compare August 18, 2026 03:00
@github-actions github-actions Bot removed the approved Indicates a PR has been approved by one committer. label Aug 18, 2026
@924060929

Copy link
Copy Markdown
Contributor

/review

@924060929

Copy link
Copy Markdown
Contributor

run buildall

@github-actions github-actions Bot added the approved Indicates a PR has been approved by one committer. label Aug 18, 2026
@github-actions

Copy link
Copy Markdown
Contributor

PR approved by at least one committer and no changes requested.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

Request changes. The patch correctly disables early SEARCH Top-K for explicit residual predicates, preserves score materialization with the existing zero sentinel, and now reports limit/offset overflow with the same exception as ordinary TopN. However, the existing P1 visibility thread remains unresolved.

The hidden delete-sign predicate protects the UNIQUE_KEYS/MOW example discussed there, but not DUP_KEYS legacy delete predicates. A lone top-level SEARCH on a DUP_KEYS table still retains a positive score limit; BE evaluates SEARCH with that limit before applying the storage delete predicate. A deleted highest scorer can therefore consume Top-1 and be removed afterward, leaving the live runner-up unavailable. I replied to the existing thread with the concrete reduced plan and FE/BE ordering evidence. No new inline thread is added here because that would duplicate the existing P1.

Review checkpoints

  • Goal, correctness, and proof: The explicit extra-predicate cases are fixed, but the retained lone-SEARCH branch still has a reachable wrong-result path on supported DUP_KEYS deletes, so the overall correctness goal is incomplete. The changed result cases are deterministic and exercise explicit residual predicates, but no deleted-highest-score regression covers this remaining path.
  • Scope and maintainability: The production change is small and readable. The helper's normalized-expression truth table is conservative for multiple, residual, nested, and negated SEARCH predicates; its unsafe assumption is that the logical conjunct set represents every later semantic mask.
  • Concurrency and lifecycle: No new shared mutable state, locking, thread lifecycle, resource lifecycle, or static initialization behavior is introduced.
  • Configuration, compatibility, and propagation: No configuration, persisted state, storage format, or protocol field changes. FE reuses the existing score-limit field, and 0 already selects BE's full-doc-set path, so mixed-path propagation is intact.
  • Parallel and special paths: SEARCH versus MATCH-only scoring, extracted min-score predicates, multiple/nested SEARCH, key ranges, parallel scanner row ranges, UNIQUE_KEYS/MOW delete-sign and bitmap handling, and DUP_KEYS legacy delete predicates were traced. Parallel row ranges are exhaustive work partitions; the blocking path is the semantic legacy-delete mask applied after SEARCH collection.
  • Error handling and observability: Overflow now matches LogicalTopNToPhysicalTopN exactly and the new exception regression distinguishes it from the prior unrelated score-usage error. EXPLAIN exposes SCORE SORT LIMIT, but the existing positive-versus-zero and recursive-classifier test-oracle gaps remain covered by the prior P2 threads.
  • Transactions, persistence, and writes: This PR does not change write or persistence code. It does affect read visibility after a committed DELETE, which is the blocking correctness issue.
  • Performance: Full collection for limit 0 is more expensive, but it is the required correctness fallback until all later semantic masks are incorporated before Top-K selection.
  • Tests and CI: The 12 changed ordered result cases and overflow case are internally consistent. Per the review-only instructions, no local build or test was run. At submission time CheckStyle and the lightweight GitHub checks pass; Doris compile and FE UT are still pending.
  • User focus: No additional review focus was provided.

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 17492 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit a96afba8eea0ae305b0404caf34d1bb6ee3601f6, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17600	3185	3190	3185
q2	1893	245	152	152
q3	10446	897	528	528
q4	4678	248	210	210
q5	7664	593	389	389
q6	143	115	92	92
q7	533	509	386	386
q8	9247	888	861	861
q9	3491	2420	2409	2409
q10	6521	859	716	716
q11	459	257	254	254
q12	673	398	331	331
q13	17894	1529	1177	1177
q14	161	154	140	140
q15	q16	430	409	371	371
q17	809	748	840	748
q18	3145	2272	2279	2272
q19	1139	937	855	855
q20	660	544	501	501
q21	5282	1685	1933	1685
q22	333	258	230	230
Total cold run time: 93201 ms
Total hot run time: 17492 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	3542	3490	3452	3452
q2	226	220	157	157
q3	2232	2345	2118	2118
q4	1221	1181	905	905
q5	2222	2144	2128	2128
q6	174	125	86	86
q7	1006	974	868	868
q8	1637	1458	1447	1447
q9	3189	3160	3144	3144
q10	1884	1832	1666	1666
q11	367	277	259	259
q12	464	432	339	339
q13	1496	1575	1180	1180
q14	169	174	159	159
q15	q16	395	401	353	353
q17	1084	1060	1060	1060
q18	5048	4419	4791	4419
q19	878	846	849	846
q20	981	945	838	838
q21	3809	3099	3148	3099
q22	398	348	312	312
Total cold run time: 32422 ms
Total hot run time: 28835 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 83904 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit a96afba8eea0ae305b0404caf34d1bb6ee3601f6, data reload: false

query5	4245	431	348	348
query6	422	169	154	154
query7	4831	438	258	258
query8	302	126	116	116
query9	8671	2940	2922	2922
query10	421	259	210	210
query11	5383	1071	943	943
query12	123	74	72	72
query13	1190	460	343	343
query14	6044	2263	2130	2130
query14_1	2033	1997	1986	1986
query15	174	118	120	118
query16	3206	380	352	352
query17	879	435	376	376
query18	2101	319	233	233
query19	179	138	109	109
query20	72	66	72	66
query21	735	117	106	106
query22	5669	5362	5310	5310
query23	6764	6147	5937	5937
query23_1	6055	6093	6026	6026
query24	7310	1107	772	772
query24_1	796	790	779	779
query25	417	286	239	239
query26	1217	269	165	165
query27	2715	460	280	280
query28	4629	1500	1510	1500
query29	914	430	361	361
query30	427	183	159	159
query31	860	440	365	365
query32	108	53	56	53
query33	507	219	189	189
query34	1017	838	482	482
query35	412	403	354	354
query36	565	550	537	537
query37	122	84	73	73
query38	1066	854	839	839
query39	493	499	497	497
query39_1	473	470	468	468
query40	313	160	115	115
query41	57	55	51	51
query42	79	83	78	78
query43	256	264	224	224
query44	1064	557	576	557
query45	107	104	106	104
query46	803	861	550	550
query47	787	764	728	728
query48	331	305	237	237
query49	635	251	205	205
query50	866	342	278	278
query51	8298	8263	8102	8102
query52	73	76	70	70
query53	211	209	154	154
query54	242	173	157	157
query55	71	58	57	57
query56	260	214	216	214
query57	945	670	633	633
query58	239	203	208	203
query59	1237	1253	1108	1108
query60	283	214	219	214
query61	145	155	127	127
query62	383	213	181	181
query63	195	155	162	155
query64	2401	791	750	750
query65	1628	1597	1631	1597
query66	1751	287	235	235
query67	9791	9870	9656	9656
query68	2763	1250	819	819
query69	518	223	207	207
query70	660	596	648	596
query71	297	273	241	241
query72	2823	1740	1561	1561
query73	689	578	358	358
query74	1571	1230	1126	1126
query75	1249	1151	1040	1040
query76	1870	746	549	549
query77	261	253	215	215
query78	3908	3634	3223	3223
query79	2923	826	578	578
query80	1513	387	349	349
query81	538	200	187	187
query82	804	133	106	106
query83	326	249	243	243
query84	375	126	103	103
query85	963	454	386	386
query86	499	175	169	169
query87	1008	990	890	890
query88	2892	2148	2143	2143
query89	319	223	208	208
query90	1963	141	146	141
query91	158	142	125	125
query92	63	46	44	44
query93	1666	1241	762	762
query94	733	249	223	223
query95	627	372	406	372
query96	824	574	283	283
query97	1040	1045	1036	1036
query98	174	137	136	136
query99	487	340	313	313
Total cold run time: 187784 ms
Total hot run time: 83904 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 14.62 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit a96afba8eea0ae305b0404caf34d1bb6ee3601f6, data reload: false

query1	0.01	0.00	0.00
query2	0.19	0.03	0.03
query3	0.49	0.11	0.11
query4	1.77	0.11	0.10
query5	0.17	0.16	0.15
query6	1.49	0.70	0.71
query7	0.04	0.01	0.00
query8	0.05	0.03	0.02
query9	0.50	0.21	0.22
query10	0.35	0.34	0.36
query11	0.23	0.12	0.11
query12	0.20	0.12	0.12
query13	0.31	0.30	0.33
query14	0.47	0.45	0.45
query15	0.40	0.37	0.35
query16	0.23	0.22	0.23
query17	0.67	0.68	0.70
query18	0.18	0.17	0.17
query19	1.53	1.20	1.13
query20	0.02	0.01	0.01
query21	16.46	0.16	0.12
query22	4.72	0.04	0.04
query23	16.86	0.25	0.11
query24	12.04	0.31	0.26
query25	0.12	0.04	0.04
query26	0.73	0.16	0.12
query27	0.04	0.03	0.03
query28	3.64	0.48	0.27
query29	12.86	3.20	2.55
query30	0.25	0.12	0.12
query31	2.81	0.37	0.17
query32	3.65	0.32	0.23
query33	1.36	1.39	1.48
query34	15.39	2.22	1.73
query35	1.75	1.72	1.72
query36	0.46	0.30	0.29
query37	0.15	0.04	0.04
query38	0.05	0.03	0.03
query39	0.08	0.02	0.02
query40	0.16	0.09	0.07
query41	0.22	0.02	0.03
query42	0.22	0.02	0.02
query43	0.04	0.03	0.03
Total cold run time: 103.36 s
Total hot run time: 14.62 s

@hello-stephen

Copy link
Copy Markdown
Contributor

FE Regression Coverage Report

Increment line coverage 29.27% (12/41) 🎉
Increment coverage report
Complete coverage report

@hello-stephen

Copy link
Copy Markdown
Contributor

FE UT Coverage Report

Increment line coverage 61.54% (8/13) 🎉
Increment coverage report
Complete coverage report

@airborne12 airborne12 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@airborne12
airborne12 merged commit 2b6a45e into apache:master Aug 21, 2026
35 of 36 checks passed
airborne12 added a commit that referenced this pull request Sep 1, 2026
…5821 (#67327)

### What problem does this PR solve?

Issue Number: N/A

Related PR: #65821 (master), picked from commit
2b6a45e

Problem Summary:

Backport of #65821 to branch-4.1. Search score TopN pushdown may return
incorrect results when the search predicate is combined with additional
predicates, because the pushed TopN limit can be applied before the
remaining predicates are evaluated. This change disables the pushed
search TopN limit in those cases while preserving the virtual score
column pushdown, and adds regression coverage (search + equality / range
/ match / score range / multiple search predicates, plus limit+offset
overflow).

**Hunk audit (source diff → this PR):**

| Source hunk | Status |
|---|---|
| `PushDownScoreTopNIntoOlapScan.java` `@@ -194,17 +194,22 @@` (overflow
guard rework + pushedScoreLimit) | **Adapted ×2**: ① branch-4.1 never
had the #64633 overflow-guard block, so the hunk's removed lines have no
counterpart here; ② `Utils.addOverflows` does not exist on 4.1 (#64633
not backported) — inlined the equivalent check `topN.getLimit() >
Long.MAX_VALUE - topN.getOffset()` (identical to the master helper's
implementation). |
| `PushDownScoreTopNIntoOlapScan.java` `@@ -243,6 +248,19 @@`
(`shouldDisableSearchTopN` helper) | Ported |
| `test_search_score_topn_predicates.out` (new) | Ported (verbatim) |
| `test_search_score_topn_predicates.groovy` (new) | **Adapted**:
dropped `set enable_segment_limit_pushdown = true` — the variable comes
from #62222 which is not on 4.1; it defaults to true on master and only
controls a BE-side segment limit optimization, unrelated to this
FE-plan-level fix. |

**Local verification on this branch:** full ASAN BE+FE build green;
`run-regression-test.sh -d inverted_index_p0 -s
test_search_score_topn_predicates` → 1 suite, 0 failed against a local
1FE+1BE cluster built from this PR. No FE UT exists for this rule on 4.1
and the source PR added none (its coverage is the regression suite
above).

### Release note

None

### Check List (For Author)

- Test <!-- At least one of them must be included. -->
    - [x] Regression test
    - [ ] Unit Test
    - [ ] Manual test (add detailed scripts or steps below)
    - [ ] No need to test or manual test. Explain why:
- [ ] This is a refactor/code format and no logic has been changed.
        - [ ] Previous test can cover this change.
        - [ ] No code files have been changed.
        - [ ] Other reason <!-- Add your reason?  -->

- Behavior changed:
    - [x] No.
    - [ ] Yes. <!-- Explain the behavior change -->

- Does this need documentation?
    - [x] No.
- [ ] Yes. <!-- Add document PR link here. eg:
apache/doris-website#1214 -->

### Check List (For Reviewer who merge this PR)

- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label <!-- Add branch pick label that this PR
should merge into -->

Co-authored-by: liangj777 <106017102+LIANG751234313@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

approved Indicates a PR has been approved by one committer. dev/4.1.x dev/4.1.x-conflict reviewed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants