Skip to content

Consider using a separate Lucene term for each hash table #293

Description

@alexklibisz

The LSH models currently encode the hash table in the hash and index them all under the same term.

For example, if the hash value for the 99th table is 42, the value that's stored in Lucene is actually something like "99.42" (more efficiently encoded, but that's the gist of it.

Another way to index the hash values would be to have a separate Lucene term for each hash table. So the 99 would be part of the term name, not the actual value. In theory this should save space, but it would require some experimentation to answer:

  • is it a meaningful amount of space
  • how does it affect query performance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    algorithmsAdding and improving ANN algorithms

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions