fix(graph-dac): push indexed equality filters into the native search query - #1314
Open
likhithThammegowda wants to merge 1 commit into
Open
likhithThammegowda wants to merge 1 commit into
likhithThammegowda wants to merge 1 commit into
Conversation
Contributor
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
likhithThammegowda
force-pushed
the
spark-v1.0.2.1-work
branch
from
September 30, 2026 09:28
2bc0cec to
36ddf19
Compare
…query
executeNativeSearch only added IL_SYS_NODE_TYPE and IL_FUNC_OBJECT_TYPE to the
JanusGraph query when the caller had set them via SearchCriteria.setNodeType() /
setObjectType(). Callers that carry the same conditions as metadata filters --
FrameworkValidator.getMasterCategoryNodes and getCategoryTermsFromDB, and
PropAsEdgeValidator -- set neither, so the base query stayed has("graphId", ...)
alone. Every node in a graph carries the same graphId, so that predicate has no
selectivity and the query returned the entire graph. matchesMetadata then filtered
in memory, and checkFilter -> vertex.property() lazy-loaded each vertex from the
storage backend one round trip at a time.
On our staging environment this made content create block until the 30s actor ask
timeout and return 500 -- while the node itself was written in 1.7s, so the content
was created and a client retry duplicated it. Reads by identifier were unaffected:
they take the extractIdsFromMetadata fast path. Setting
master.category.validation.enabled=No, which skips the master-category lookup, took
create from a 30146ms failure to a 3777ms success, isolating this as the cause. The
composite indexes were present and ENABLED throughout; the query simply never used
them.
Indexable equality predicates carried in the metadata are now added to the graph
query. Only conditions that must hold for every match are pushed:
- criteria combined with OR are skipped, since an OR branch need not be true and
pushing it would exclude valid matches
- only OP_EQUAL is pushed; a composite index cannot serve !=, range or list
predicates, so status != Retired continues to be applied in memory
- only keys the graph schema indexes, with a String value, are pushed
This narrows the candidate set only -- the in-memory filter still runs over the
result, so the returned nodes are unchanged. The key list is configurable via
graph.native_search.indexed_keys and an empty list restores the previous behaviour.
Tests cover the regressed master-category query shape, OR criteria being left
in memory, indexed and non-indexed keys filtering identically, the two combined,
and the unfiltered case.
likhithThammegowda
force-pushed
the
spark-v1.0.2.1-work
branch
from
September 30, 2026 13:03
1a4462e to
f290e5f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Content create (
POST /content/v3/create) blocks for 30 seconds and returns 500, although the node is written in 1.7s. The caller never learns the content was created, so a client retry creates a duplicate.This PR now contains one commit. The other three it used to carry (
99a6a8fd,b35467dd,2bc0cece) were merged via #1317 on 28 September, so the branch has been rebased ontospark-v1.0.2.1to leave only what is still new.Problem
Seen in the creation portal when creating a course, and when saving a resource in the Course Builder: the spinner ran for 30 seconds and failed, but the course had been created, so retrying produced duplicates in My Courses.
The 500 is the Pekko ask on
contentActortiming out, not the write failing. Reads by identifier are unaffected: they take theextractIdsFromMetadatafast path.Cause
executeNativeSearchaddsIL_SYS_NODE_TYPEandIL_FUNC_OBJECT_TYPEto the JanusGraph query only when the caller sets them viaSearchCriteria.setNodeType()/setObjectType(). Callers that carry the same conditions as metadata filters set neither:FrameworkValidator.getMasterCategoryNodesFrameworkValidator.getCategoryTermsFromDBPropAsEdgeValidatorSo the base query stays
has("graphId", ...)alone. Every node in a graph has the samegraphId, so that predicate has no selectivity and the query returns the entire graph.matchesMetadatathen filters in memory, andcheckFilter→vertex.property()lazy-loads each vertex from the storage backend one round trip at a time.Isolation: setting
master.category.validation.enabled=No, which skips the master-category lookup, took create from a 30146 ms failure to a 3777 ms success. The composite indexes were present and ENABLED throughout; the query simply never used them.Fix
Indexable equality predicates carried in the metadata are now added to the graph query. Only conditions that must hold for every match are pushed:
ORare skipped: an OR branch need not be true, and pushing it would exclude valid matchesOP_EQUALis pushed: a composite index cannot serve!=, range or list predicates, sostatus != Retiredstill applies in memoryThis narrows the candidate set only. The in-memory filter still runs over the result, so the returned nodes are unchanged. The key list is configurable via
graph.native_search.indexed_keys; an empty list restores the previous behaviour.ontology-engine/graph-dac-api/.../SearchAsyncOperations.javaexecuteNativeSearch(+60)ontology-engine/graph-dac-api/.../SearchAsyncOperationsPushDownTest.javaType of change
How Has This Been Tested?
POST /content/v3/createreturned 200 in 3546 ms with framework validation still enabled, where it previously failed after 30146 ms. The course appeared once, with no duplicate.SearchAsyncOperationsPushDownTestis added and covers the regressed master-category query shape, OR criteria staying in memory, indexed and non-indexed keys filtering identically, both combined, and the unfiltered case. It has not been run locally; please let CI confirm.The rebased branch was checked against the image deployed on staging:
spark-v1.0.2.1plus this commit produces identical source to what is running there.Test Configuration:
spark-v1.0.2.1)Checklist: