Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 1 addition & 4 deletions .husky/pre-commit
Original file line number Diff line number Diff line change
@@ -1,4 +1 @@
#!/bin/sh
. "$(dirname "$0")/_/husky.sh"

yarn lint-staged
pnpm lint-staged
4 changes: 4 additions & 0 deletions .oxlintrc.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
{
"$schema": "./node_modules/oxlint/configuration_schema.json",
"ignorePatterns": ["tantivy-py/**"]
}
4 changes: 3 additions & 1 deletion .prettierignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,4 +5,6 @@ package-template.wasi-browser.js
package-template.wasi.cjs
wasi-worker-browser.mjs
wasi-worker.mjs
.yarnrc.yml
.yarnrc.yml
tantivy-py
pnpm-lock.yaml
2 changes: 1 addition & 1 deletion .taplo.toml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
exclude = ["node_modules/**/*.toml"]
exclude = ["node_modules/**/*.toml", "tantivy-py/**/*.toml"]

# https://taplo.tamasfe.dev/configuration/formatter-options.html
[formatting]
Expand Down
4 changes: 3 additions & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -10,14 +10,16 @@ crate-type = ["cdylib"]

[dependencies]
chrono = { version = "0.4", features = ["serde"] }
futures = "0.3"
napi = { version = "3.12", default-features = false, features = [
"napi8",
"serde-json",
] }
napi-derive = "3.6"
serde = { version = "1.0", features = ["derive"] }
serde_json = { version = "1.0", default-features = false, features = ["alloc"] }
tantivy = "0.25.0"
tantivy = "0.26.2"
tantivy-common = "0.11"

[build-dependencies]
napi-build = "2.4"
Expand Down
81 changes: 57 additions & 24 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,6 @@ Node.js bindings for [Tantivy](https://github.com/quickwit-oss/tantivy), the ful

This project is a Node.js port of [tantivy-py](https://github.com/quickwit-inc/tantivy-py), providing JavaScript/TypeScript bindings for the Tantivy search engine. The implementation closely follows the Python API to maintain consistency across language bindings.



# Installation

The bindings can be installed using npm:
Expand Down Expand Up @@ -60,8 +58,11 @@ This Node.js binding provides access to most of Tantivy's functionality:
- **Snippet generation** for search result highlighting
- **Query explanation** for debugging relevance scoring
- **Multiple field types**: text, integers, floats, dates, facets
- **Flexible tokenization** and text analysis
- **JSON document support**
- **Flexible tokenization** and text analysis, including fast-field tokenizers
- **JSON document support**, with optional dot-path expansion
- **Aggregations**, including cardinality
- **Fast-field reads** (`fastFieldValues`) and prefix term lookup (`termsWithPrefix`)
- **Schema-free query parsing** via `parseQuery` / `parseQueryLenient`

## API Compatibility

Expand Down Expand Up @@ -115,18 +116,21 @@ The Node.js implementation differs from the Python version in several ways:

#### 🔴 Critical Validation Issues

##### Numeric Field Validation (Too Lenient)
##### Multi-valued Single Fields (Too Lenient)

**Current behavior**: Node.js version accepts invalid values that Python rejects
**TODO**: Implement strict validation to match Python behavior
**Current behavior**: an array is accepted wherever a single value is expected
**TODO**: decide whether to match Python, which rejects it

```javascript
// ❌ These currently PASS in Node.js but should FAIL:
Document.fromDict({ unsigned: -50 }, schema) // Should reject negative for unsigned
Document.fromDict({ signed: 50.4 }, schema) // Should reject float for integer
Document.fromDict({ unsigned: [1000, -50] }, schema) // Should reject arrays for single fields
// ❌ This currently PASSES in Node.js but should FAIL:
Document.fromDict({ unsigned: [1000, 50] }, schema) // Should reject arrays for single fields
```

Numeric values themselves are validated as in tantivy-py: a negative value for an
unsigned field, or a fractional value for an integer field, throws rather than
being silently coerced. Values beyond `Number.MAX_SAFE_INTEGER` can be passed as
a `BigInt`.

##### Bytes Field Validation (Too Restrictive)

**Current behavior**: Only accepts Buffer objects
Expand Down Expand Up @@ -157,26 +161,55 @@ Document.fromDict({ json: 123 }, schema) // Should reject numbers
Document.fromDict({ json: 'hello' }, schema) // Should reject strings
```

#### 🟠 Error Handling Differences
#### 🔵 Type System Differences

##### Fast Field Configuration
##### Date Handling

**Current**: Throws exception when field not configured as fast
**Python**: Returns empty results
**TODO**: Decide on consistent error handling approach
**Current**: Dates cross the boundary as millisecond timestamps. `Document.fromDict()`
and the query builders also accept a JavaScript `Date` or an ISO 8601 string, and
`toDict()` returns milliseconds since the epoch.
**Python**: Uses `datetime` objects with nanosecond storage.

##### Query Parser Errors
Sub-millisecond precision is therefore not representable from JavaScript.

**Current**: Different error message formats
**TODO**: Align error messages with Python version
#### 🔵 Resource Management

#### 🔵 Type System Differences
`IndexWriter` has no equivalent of tantivy-py's `with index.writer() as writer:`
block, because JavaScript has no deterministic destructor. Call
`writer.waitMergingThreads()` when you are done with a writer — it commits nothing
on its own, so commit first, and it is what releases the directory lock:

##### Date Handling
```javascript
const writer = index.writer()
try {
writer.addDocument(doc)
writer.commit()
} finally {
writer.waitMergingThreads()
}
```

#### 🔵 Concurrency

tantivy-py releases the GIL around the blocking calls (`addDocument`, `commit`,
`search`, …). Node has no GIL, but these methods run synchronously on the main
thread and are not offloaded to the libuv thread pool.

#### 🔵 Naming

Two shapes differ from tantivy-py because napi-rs models them differently:

- Static constructors live on a companion class: `FilterStatic.lowercase()` and
`TokenizerStatic.simple()` rather than `Filter.lowercase()` / `Tokenizer.simple()`.
- `DocAddress`, `SearchResult` and `SearchHit` are plain objects rather than
classes, so they have no constructors or getters — read `hit.docAddress`,
`hit.score` and `hit.order` directly.

`FieldType` uses tantivy-py's variant names (`Text`, `Unsigned`, `Integer`,
`Float`, `Boolean`, `Json`), not tantivy's Rust `Type` names.

**Current**: Uses getTime() timestamps
**Python**: Uses datetime objects
**TODO**: Consider more intuitive date API
Python's pickle hooks (`__reduce__`, `__getnewargs__`) have no counterpart;
`Schema` exposes `toJSON()` / `Schema.fromJson()` instead.

## Architecture

Expand Down
72 changes: 72 additions & 0 deletions __test__/fixtures.ts
Original file line number Diff line number Diff line change
Expand Up @@ -214,3 +214,75 @@ export const createSpanishIndex = () => {
index.reload()
return index
}

export const schemaOrderFastFields = () => {
return new SchemaBuilder()
.addTextField('title', { stored: true })
.addUnsignedField('u64_field', { fast: true })
.addIntegerField('i64_field', { fast: true })
.addFloatField('f64_field', { fast: true })
.addBooleanField('bool_field', { fast: true })
.addDateField('date_field', { fast: true })
.addTextField('str_field', { fast: true })
.build()
}

export const createIndexWithOrderFastFields = (dir?: string) => {
const index = new Index(schemaOrderFastFields(), dir)
const writer = index.writer(15_000_000, 1)

const low = new Document()
low.addText('title', 'low title')
low.addUnsigned('u64_field', 0)
low.addInteger('i64_field', -10)
low.addFloat('f64_field', 1.5)
low.addBoolean('bool_field', false)
low.addDate('date_field', Date.UTC(2024, 0, 1))
low.addText('str_field', 'apple')
writer.addDocument(low)

const high = new Document()
high.addText('title', 'high title')
high.addUnsigned('u64_field', 2)
high.addInteger('i64_field', 5)
high.addFloat('f64_field', 3.14)
high.addBoolean('bool_field', true)
high.addDate('date_field', Date.UTC(2026, 0, 1))
high.addText('str_field', 'cherry')
writer.addDocument(high)

writer.commit()
writer.waitMergingThreads()
index.reload()
return index
}

export const createIndexWithEmptyFastField = () => {
const indexSchema = new SchemaBuilder()
.addTextField('title', { fast: true, stored: true })
.addTextField('body', { fast: true })
.build()

const index = new Index(indexSchema)
const writer = index.writer(15_000_000, 1)

writer.addDocument(
Document.fromDict({
title:
'A record of the statesmanship and political achievements of Gen. Winfield Scott Hancock, ' +
'regular Democratic nominee for president of the United States',
}),
)
writer.addDocument(Document.fromDict({ title: 'Political Achievements of the Earl of Dalkeith' }))
writer.addDocument(
Document.fromDict({
title: 'The Old Man and the Sea',
body: 'He was an old man who fished alone inthe Gulf Stream and he had gone eighty-four days now without taking a fish.',
}),
)

writer.commit()
writer.waitMergingThreads()
index.reload()
return index
}
Loading
Loading