Skip to content

Add clarification that query caches are part of the JVM heap allocation - #1673

Open
renetapopova wants to merge 9 commits into
neo4j:devfrom
renetapopova:dev-query-caches
Open

renetapopova wants to merge 9 commits into
neo4j:devfrom
renetapopova:dev-query-caches

Conversation

@renetapopova

@renetapopova renetapopova commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

@renetapopova renetapopova changed the title Dev query caches Add clarification that query caches are part of the JVM heap allocation Sep 29, 2026
@renetapopova

renetapopova commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Hey @JPryce-Aklundh, could you please take a look at this? I am not sure why it shows a conflict.
When I tried to rebase the branch, I got:

➜  docs-cypher git:(dev-query-caches) git rebase dev
Current branch dev-query-caches is up to date.
error: refs/remotes/upstream/add-syntax-snippets-2026.09 does not point to a valid object!

Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated
Comment thread modules/ROOT/pages/planning-and-tuning/caching.adoc Outdated

Query caches may consume a lot of memory, especially when running many active databases.
To tackle this and improve predictability on memory consumption, you can configure the DBMS to use only one set of caches for all databases.
By default, the set of query caches is *per database*, which means that the memory allocated for query caches for each database is part of the JVM heap of that database.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"which means that the memory allocated for query caches for each database is part of the JVM heap of that database."

There is not really anything like "a part of the JVM heap of that database". I think it may be unnecessary to mention this part about the heap, or alternatively just have a generic note "query caches are allocated on the JVM heap" somewhere.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The request was to explicitly state that query caches are part of the JVM heap, and given that they are per database, I assumed that the memory allocated for query caches for each database is part of the JVM heap for that database.
If they are not part of the JVM heap for each of the databases, that means that the diagram here is not correct https://github.com/neo4j/docs-operations/pull/3330/changes.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, as a reference to the diagram I guess it could work conceptually.

@renetapopova renetapopova Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is the diagram correct or should I replace it with something like
neo4j-memory-management 3 (1)

For more information, see link:{neo4j-docs-base-uri}/operations-manual/current/performance/memory-configuration/[Operations Manual -> Memory configuration].

Query caches can use significant memory, especially when running many active databases.
The amount of memory consumed by query caches depends on the available heap size, the nature of the queries, and the number of unique queries executed. +

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"depends on the available heap size" is incorrect, the amount of memory consumed never depends on the available heap size.
What I meant when we chatted about this was that if query cache memory usage is even a problem to worry about depends on the available heap size. E.g. if you are running the smallest Aura Free instance the heap usage of the query cache could be a problem (if some of the other conditions also apply), but if you are running a larger instance with a larger heap you most likely never have to worry about query cache size since it is such a small fraction of the total heap size.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am not sure I understand you. "depends on the available heap size, the nature of the queries, and the number of unique queries executed." means exactly what you explain. For example, if you don't have enough heap size, and you run many unique queries, you might have a problem.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll try to rewrite it.

Comment on lines +156 to +157
In many cases, executing the queries itself consumes much more memory, as heap usage can scale with the amount of data.
This is particularly evident in multi-tenant scenarios where numerous databases run workloads of programmatically generated queries without appropriate parameterization.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the ordering of these two sentences introduces confusion. The second statement is referring to the earlier section of the conditions when query cache memory usage can become a problem and should come first.
The first statement is a general principle and a reason why query caches is not the first place to consider when tuning memory usage: heap usage of executing queries can scale with the amount of data, but the heap usage of query caches cannot.

The amount of memory consumed by query caches depends on the available heap size, the nature of the queries, and the number of unique queries executed. +
In many cases, executing the queries itself consumes much more memory, as heap usage can scale with the amount of data.
This is particularly evident in multi-tenant scenarios where numerous databases run workloads of programmatically generated queries without appropriate parameterization.
To tackle this and improve predictability of memory consumption, you can configure the DBMS to use a single set of caches for all databases.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We could clarify this with "improve predictability of memory consumption in multi-tenant use-cases ...", as with a single database or even in most use-cases with multiple databases used for data federation (composite database) this setting doesn't do anything useful.
The difference between a federated and a multi-tenancy usecase is that the latter has many copies of the "same" database (same schema/data model) running the same queries.

By default, the set of query caches is *per database*, which means that the memory allocated for query caches for each database is part of the JVM heap of that database.
For more information, see link:{neo4j-docs-base-uri}/operations-manual/current/performance/memory-configuration/[Operations Manual -> Memory configuration].

Query caches can use significant memory, especially when running many active databases.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe it is more helpful to state that "Query caches can use significant memory in certain conditions" and then go on to list the conditions? This phrasing with "especially" suggests to me that query cache memory usage is something I should be worried about when in many use cases I don't need to.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This sentence existed before. I just changed "a lot of" with "significant".

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes I know. It was not about your specific change but mostly that I don't think what existed before was very helpful.


Query caches may consume a lot of memory, especially when running many active databases.
To tackle this and improve predictability on memory consumption, you can configure the DBMS to use only one set of caches for all databases.
By default, the set of query caches is *per database*, which means that the memory allocated for query caches for each database is part of the JVM heap of that database.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
By default, the set of query caches is *per database*, which means that the memory allocated for query caches for each database is part of the JVM heap of that database.
By default, the set of query caches is *per database*, and their memory is allocated on the JVM heap.

@neo4j-docops-agent

Copy link
Copy Markdown
Collaborator

This PR includes documentation updates
View the updated docs at https://neo4j-docs-cypher-1673.surge.sh

Updated pages:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants