Skip to content

BGE Reranker V2 M3

BAAI/bge-reranker-v2-m3

BGE Reranker V2 M3 by Reranker.

Context
8.2K tokens
Endpoint
Get API KeyCompare

Pricing

Input$0.14 / 1M
Output$0.14 / 1M
Cache Write (5m)Not applicable
Cache Write (1h)Not applicable
Cache ReadNot applicable
Web Search$0 / 1M

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint:
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="BAAI/bge-reranker-v2-m3",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="BAAI/bge-reranker-v2-m3",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common4 params
modelquerydocumentstop_n
Extended2 params
return_documentsmax_chunks_per_doc

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

BAAI/bge-reranker-v2-m3

Decision guidance

Fit the request contract before production.

Use this model when

  • The workload fits the text or chat tasks shown in this model's current catalog record.
  • The input fits within the published 8.2K-token context record, with output and system overhead budgeted separately.

Check before production

  • Confirm the production client uses one of the listed request surfaces: /v1/rerank.
  • Estimate a representative request from the current Pay As You Go price fields instead of extrapolating from a tiny prompt.
  • Check the observed-availability card and run your own timeout, retry, and fallback test before relying on the model.

Compare with Other Models

See how this model compares to others from the same provider.