An OpenAI-compatible API, and where it saves money

Point any OpenAI client at https://aiyellowjacket.com/v1 and it works. Here is exactly what each endpoint does, including the one that is not cheaper than calling the model yourself.

Endpoints

EndpointWhat it does
POST /v1/chat/completionsOpenAI-style chat: send messages, get a chat.completion back with choices and usage. Answered by Claude Opus 5.5.
GET /v1/modelsLists one model, colony.
POST /v1/jobsThe document engine: a task in plain words plus documents (text or base64 images), links or CSV tables. A top model plans it, cheap models read, every answer is checked.
POST /v1/debateTwo or three of Claude, ChatGPT, Gemini and DeepSeek answer the same question, review each other's answers, and a model that took no part gives the verdict.
POST /v1/askA plain-English question about a database or spreadsheet you connected on the work desk: the AI writes the SQL, the database does the arithmetic.

The base URL is https://aiyellowjacket.com/v1. Every request carries Authorization: Bearer col_..., a key you make under Settings, For developers, on the work desk.

Point an OpenAI client at it

With the OpenAI Python library, change the base URL and the key:

from openai import OpenAI

client = OpenAI(base_url="https://aiyellowjacket.com/v1", api_key="col_YOUR_KEY")
reply = client.chat.completions.create(
    model="colony",
    messages=[{"role": "user", "content": "Draft a polite reminder for an overdue invoice."}],
)
print(reply.choices[0].message.content)

Or with curl:

curl https://aiyellowjacket.com/v1/chat/completions \
  -H "Authorization: Bearer col_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "colony",
       "messages": [{"role": "user", "content": "Write three subject lines for a payment reminder."}]}'

What /v1/chat/completions does, and doesn't

  • The model: every reply comes from Claude Opus 5.5. The model you send is not used to choose another one.
  • Privacy: unless the workspace has switched it off, names, emails, phone numbers, addresses and account numbers are swapped for placeholders before the request goes out and put back in the reply, and the request goes only to providers that keep no copy of it.
  • Shape: reads messages (text content; image parts are ignored) and max_tokens (default 16,000). Replies come back whole: no streaming, no tools or function calls. The response has choices, usage and a colony field with the model cost and what you were charged.
  • Price: our AI cost plus 25%, so a reply costs about 25% more than calling Claude Opus with your own account. What you get for that is the privacy step, zero-retention routing and one bill.

Where the savings are: /v1/jobs

Document work is where the engine pays off. Send a task in plain words and the documents; a top model plans the job once, cheap models read, code checks every answer, and the result comes back as JSON. You pay 35% of what the cheapest top AI model would charge for the same job, or our AI cost plus 25% when that is higher, and a planned job never costs more than the cheapest top model would. On 80 real contracts the AI cost was $0.50 against $4.41 for Claude Opus alone, at the same accuracy.

import requests

job = requests.post(
    "https://aiyellowjacket.com/v1/jobs",
    headers={"Authorization": "Bearer col_YOUR_KEY"},
    json={
        "task": "For each invoice: supplier, invoice number, date, total. Flag any whose lines don't add up.",
        "documents": [{"id": "inv-001", "text": open("inv-001.txt").read()},
                      {"id": "inv-002", "text": open("inv-002.txt").read()}],
        "wait": True,
    },
    timeout=960,
).json()
print(job["status"], job.get("answer"))

Documents are {"id", "text"} or {"id", "image_base64", "mime"}; links takes public URLs and tables takes CSV text. With wait the call returns when the job finishes (up to 15 minutes); otherwise poll GET /api/jobs/{id} with the same key, or pass a callback_url and receive a signed job.finished webhook.

Keys and billing

Make a key under Settings, For developers, on the work desk. Work is paid from prepaid credit; a request that the credit can't cover is answered with HTTP 402. To run the models on your own OpenRouter account, add the header X-OpenRouter-Key: OpenRouter bills you for the models and we charge 20% of the model cost.

Questions people ask

Is this cheaper than the OpenAI or Anthropic API?

For chat, no: /v1/chat/completions is Claude Opus at our AI cost plus 25%. The savings are in /v1/jobs, where cheap models do the reading under checks and you pay a share of what a top model would charge.

Which model answers /v1/chat/completions?

Claude Opus 5.5. /v1/models lists one model, colony.

Does it support streaming or function calling?

Not yet. Replies come back whole, and tools or function definitions are not used.

What is the base URL?

https://aiyellowjacket.com/v1, with Authorization: Bearer and a col_ key from Settings, For developers.

Take the sting out of your AI bill

Run a job