Mistral Embed is sold per token (currently about $0.10 per 1M input tokens, 1024-dim,
8K context). At that rate a RAG agent's embed call is a sub-cent event: 500 input
tokens is roughly 0.005 cents, 1,000 tokens about a tenth of a cent. That is exactly
the granularity where a fixed card fee stops making sense - a single $0.30 card charge
(2.9% + floor) equals about 3,000 one-thousand-token embed calls.
The offer: an optional feeless settlement rail beside the existing per-token meter, so
an agent can settle an embedding call live with no card, no pre-fund, no per-transaction
fee floor. Shape:
- For the SDK: an optional settlement flag on the client, default unset = unchanged:
from mistralai import Mistral
client = Mistral(api_key=..., settlement="nano") # optional feeless rail
- For billing/usage: accept a feeless rail as a settlement currency for metered usage,
so per-request micropayments clear instantly with no fee floor, beside the existing
card/usage billing. The unit meter is already there - a feeless final rail completes
the "pay for what you embed" pitch for autonomous agents whose per-call cost is
smaller than any card floor.
Why a feeless rail, not just micro-billing: the x402 HTTP-402 payment spec
(https://x402.org, Nano binding https://x402nano.org) and a reference feeless402 helper
(https://github.com/michardnicolas/feeless402) already turn a metered request into an
atomic pay-for-use call with no minimum and no pre-fund. A thin, opt-in client path -
not a new payment system. This is the same reasoning behind x402's own support for
per-request agent payments.
Alternatives considered:
- USDC/x402 on Base: works, but the on-chain gas floor exceeds a sub-cent embed.
- Card micro-billing: $0.30 fixed fee is a floor that dwarfs a ramping agent's embeds.
- Do nothing: agents batch embed calls or self-host open models; the per-call value stays unrealised.
Happy to shape this to your CONTRIBUTING and open a PR with the client-side settlement
option if you would consider it. I can share the small tested cost model I built from
your published rate.
Mistral Embed is sold per token (currently about $0.10 per 1M input tokens, 1024-dim,
8K context). At that rate a RAG agent's embed call is a sub-cent event: 500 input
tokens is roughly 0.005 cents, 1,000 tokens about a tenth of a cent. That is exactly
the granularity where a fixed card fee stops making sense - a single $0.30 card charge
(2.9% + floor) equals about 3,000 one-thousand-token embed calls.
The offer: an optional feeless settlement rail beside the existing per-token meter, so
an agent can settle an embedding call live with no card, no pre-fund, no per-transaction
fee floor. Shape:
from mistralai import Mistral
client = Mistral(api_key=..., settlement="nano") # optional feeless rail
so per-request micropayments clear instantly with no fee floor, beside the existing
card/usage billing. The unit meter is already there - a feeless final rail completes
the "pay for what you embed" pitch for autonomous agents whose per-call cost is
smaller than any card floor.
Why a feeless rail, not just micro-billing: the x402 HTTP-402 payment spec
(https://x402.org, Nano binding https://x402nano.org) and a reference feeless402 helper
(https://github.com/michardnicolas/feeless402) already turn a metered request into an
atomic pay-for-use call with no minimum and no pre-fund. A thin, opt-in client path -
not a new payment system. This is the same reasoning behind x402's own support for
per-request agent payments.
Alternatives considered:
Happy to shape this to your CONTRIBUTING and open a PR with the client-side settlement
option if you would consider it. I can share the small tested cost model I built from
your published rate.