Wetware API
OpenAI-compatible by interface. Human-operated by design.
Use the ordinary OpenAI SDK with a Wetware base URL, API key, and model ID. Models are people, and availability is intentionally bounded.
Quick start
Install the official OpenAI Python SDK, then point it at Wetware.
from openai import OpenAI
client = OpenAI(
api_key="sk-live-example",
base_url="http://localhost:8080/v1",
)
response = client.chat.completions.create(
model="moses",
messages=[{"role": "user", "content": "hello"}],
)
print(response.choices[0].message.content)
Models
List the human-operated models visible to your API key or retrieve one by ID.
/v1/models/v1/models/{model}Chat completions
/v1/chat/completionsWetware accepts a model ID and OpenAI-style message history. System, developer, user, assistant, and tool-result messages are presented to the operator. The operator submits either one text response or permitted function calls.
{
"model": "moses",
"messages": [
{"role": "user", "content": "hello"}
]
}
Responses
/v1/responsesUse a string or text-message history as input. Optional instructions is shown as a leading developer message. The same operator console, availability rules, capacity limit, and deadline apply.
response = client.responses.create(
model="moses",
input="What do you think about this architecture?",
instructions="Be concise and specific.",
store=False,
)
print(response.output_text)
This is a stateless subset. Supply your own history, including function or custom calls and matching result items. Direct tool namespaces and developer additional_tools catalogs are supported. Cloud-hosted tools, opaque reasoning history, background execution, response retrieval, and previous_response_id are not. JSON-object text output is available through text={"format": {"type": "json_object"}}; text-output JSON Schema is not.
Unlike OpenAI, store defaults to false; true is rejected. This is not a zero-retention setting: accepted requests and human replies still persist in Wetware's internal history.
Bounded cache/client-metadata hints are accepted and discarded, not persisted or echoed. Reasoning effort is guidance for the operator, not a hidden reasoning engine; optional encrypted-reasoning inclusion does not generate encrypted content. Compatibility with a coding harness is a defined subset, not universal feature parity.
Function tools
A human can request a function call through the same interface as a model. Your application provides definitions; the operator chooses a function and writes JSON arguments. Wetware validates and returns the call. Your application, not Wetware or the operator console, executes it and sends the result in a new request.
Chat Completions nests a definition under {"type":"function","function":{...}}. Responses places the same fields directly under {"type":"function",...}. Both accept tool_choice values auto, none, and required. A named choice is {"type":"function","function":{"name":"double_integer"}} for Chat or {"type":"function","name":"double_integer"} for Responses.
auto permits text or calls; none forbids calls; required requires at least one; a named choice requires exactly one call to that function. parallel_tool_calls=false also caps a turn at one call. With tools, parallel calls default to enabled, up to eight. This does not increase the human model's inference capacity.
Use this deterministic local example with the client above. It only allows one named arithmetic function, independently checks arguments, and performs no network, file, or shell operations.
import json
definition = {
"name": "double_integer",
"description": "Double a small integer in the caller application.",
"parameters": {
"type": "object",
"properties": {"value": {"type": "integer", "minimum": -1000, "maximum": 1000}},
"required": ["value"],
"additionalProperties": False,
},
"strict": True,
}
def run_allowed_function(name, arguments):
if name != "double_integer":
raise ValueError("Function is not allowed")
args = json.loads(arguments)
if (type(args) is not dict or set(args) != {"value"}
or type(args["value"]) is not int or not -1000 <= args["value"] <= 1000):
raise ValueError("Invalid arguments")
return json.dumps({"doubled": args["value"] * 2})
For a complete two-round Responses conversation, retain the returned calls and append one result per call_id:
tools = [{"type": "function", **definition}]
history = [{"role": "user", "content": "Use double_integer for 21, then report the result."}]
first = client.responses.create(
model="moses", input=history, tools=tools,
tool_choice={"type": "function", "name": "double_integer"},
parallel_tool_calls=False, store=False,
)
history.extend(item.model_dump(exclude_none=True) for item in first.output)
for item in first.output:
if item.type == "function_call":
result = run_allowed_function(item.name, item.arguments)
history.append({"type": "function_call_output", "call_id": item.call_id, "output": result})
final = client.responses.create(
model="moses", input=history, tools=tools, tool_choice="none", store=False,
)
print(final.output_text)
For Chat Completions, retain the assistant's tool_calls and append tool-role messages instead:
tools = [{"type": "function", "function": definition}]
history = [{"role": "user", "content": "Use double_integer for 21, then report the result."}]
first = client.chat.completions.create(
model="moses", messages=history, tools=tools,
tool_choice="required", parallel_tool_calls=False,
)
assistant = first.choices[0].message
calls = assistant.tool_calls or []
history.append({
"role": "assistant", "content": assistant.content,
"tool_calls": [call.model_dump(exclude_none=True) for call in calls],
})
for call in calls:
result = run_allowed_function(call.function.name, call.function.arguments)
history.append({"role": "tool", "tool_call_id": call.id, "content": result})
final = client.chat.completions.create(
model="moses", messages=history, tools=tools, tool_choice="none",
)
print(final.choices[0].message.content)
The human submits either text or tool calls, never both. Calls complete the inference and release capacity; the next request needs an online model and a free slot. There is no offline queue, reserved workflow, or promise that the same operator session handles the next round. Every historical call needs exactly one matching result before continuing.
Schema and size limits
Arguments must be a JSON object matching the supplied parameter schema, even with strict=false. Wetware defaults omitted strict to false on both endpoints and never rewrites schemas. Unlike upstream Responses, it does not attempt automatic strict normalization. Explicit strict=true requires additionalProperties=false and every declared property in required for each object.
Supported structural keywords are type, properties, required, boolean additionalProperties, object-schema items, enum, const, title, description, string/array/object minimum and maximum lengths or counts, and numeric bounds or multipleOf. An optional $schema must identify JSON Schema 2020-12. Bounded default and boolean encrypted annotations are accepted, but Wetware does not populate defaults or encrypt individual fields. References, composition, patterns, formats, and external resources are rejected. Parameter validation does not enable JSON Schema text output.
Limits are 64 flattened definitions, 8 calls per result, 16 KiB per description/schema, schema depth 16, and 256 schema nodes. JSON numbers in schemas, arguments, and historical arguments are limited to 128 bytes, explicit exponents from -308 to +308, and no float64 overflow. Existing aggregate request, prompt, and response byte limits also apply. Malformed definitions and unmatched history fail before notifying the operator. A rejected operator submission can be corrected before the request deadline.
Function definitions, arguments, and results are stored privately like prompts. The caller controls its function allowlist, permissions, approval requirements, and execution deduplication. Stable call IDs survive idempotent replay; do not execute an already-completed side effect again.
Custom tools and namespaces
Responses supports {"type":"custom","name":"apply_patch","format":{"type":"text"}}. The operator writes exact nonempty text, not JSON arguments. Wetware returns a custom_tool_call with input; the caller approves and executes it, then returns custom_tool_call_output with the matching call_id. Whitespace is preserved. Chat Completions remains function-only.
A custom format may provide a Lark or regex grammar of at most 16 KiB. Constraint supplied by caller; Wetware does not execute grammar. It also does not check grammar conformance. The caller must validate the completed input before execution; a patch string is not authorization to change files.
One-level namespace wrappers group directly supplied tools and retain a separate namespace and leaf name in returned calls. The console shows namespace routing guidance. Use auto or required for namespaced tools; named choices target top-level tools only. Nested namespaces, deferred loading, and hosted tools are unsupported. Developer-role additional_tools input items merge direct definitions under the same catalog and request limits.
Anthropic Messages
/v1/messages/v1/messages/count_tokensThe same human runtime is available through a bounded Messages adapter. Use a Wetware API key, not an Anthropic key. The pinned official Python SDK is anthropic==1.8.0; its base URL is the server root, without /v1.
from anthropic import Anthropic
client = Anthropic(
api_key="sk-live-example",
base_url="http://localhost:8080",
max_retries=0,
)
message = client.messages.create(
model="moses", max_tokens=1024,
messages=[{"role": "user", "content": "hello"}],
)
for block in message.content:
if block.type == "text":
print(block.text)
Direct HTTP callers provide anthropic-version: 2023-06-01 and either x-api-key or Bearer authorization, not both. Text and caller-executed function tools use standard tool_use/tool_result blocks. Thinking, cloud-hosted tools, and opaque reasoning are unsupported. Maximum tokens and effort guide the operator, not a tokenizer.
Messages usage and count_tokens are approximate byte-derived estimates, identified by X-Wetware-Token-Counting: approximate. They are not Claude tokenizer counts or billing units. Counting requires authentication and applies rate limits, but works without online presence and never sends a human request. OpenAI responses still use usage: null.
Inline images
The console can display bounded inline PNG/JPEG images, including tool results from a caller-side image viewer. Use supported data-URL image parts with OpenAI protocols or base64 image sources with Messages. No remote URL is fetched. SVG, PDF, audio, and generated-image output are unsupported.
Images are visible to the human operator and retained with inference history and backups. Body and normalized-prompt limits include base64 overhead, so large screenshots may be rejected. Operators can raise bounded deployment limits after reviewing memory and abuse exposure; the default request limit remains 256 KiB.
Streaming
Set stream=true to receive OpenAI-compatible server-sent events. The operator composes a complete response first; Wetware then serializes it. Chat Completions uses chunks followed by [DONE]. Responses uses named events, including response.output_text.delta, and ends with response.completed. Human keystrokes and internal presence events are never streamed.
Function calls stream through indexed Chat delta.tool_calls or Responses response.function_call_arguments.delta/.done events. Accumulate the complete validated call before deciding whether to execute it; partial argument fragments are not executable instructions.
Custom Responses calls use response.custom_tool_call_input.delta/.done. Messages streams use message_start, content-block events, message_delta, and message_stop, with input_json_delta for function arguments. Neither Responses nor Messages emits [DONE]. All wait for the human's full submission first.
with client.responses.stream(
model="moses", input="hello", store=False,
) as stream:
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
response = stream.get_final_response()
Availability
A model is callable only while an authenticated operator connection holds a live presence lease and capacity is free. There is no offline queue.
- Online and available
- The request is accepted and waits for the bounded human response.
- Offline
- The request fails immediately with
503 model_unavailable. - At capacity
- The request fails immediately with
429 model_at_capacity. - Response deadline
- The request fails with a provider-style timeout error. Public connections are never held indefinitely.
Errors
OpenAI endpoints use the OpenAI error shape; Messages endpoints use the Anthropic error envelope. Operator details such as browser state, location, or last-seen time are not exposed.
{
"error": {
"message": "The model `moses` is temporarily unavailable.",
"type": "service_unavailable",
"param": null,
"code": "model_unavailable"
}
}
Security and data
API keys are shown once and stored as keyed digests. Operator sessions use secure cookies and browser mutations require a CSRF token. Wetware stores accepted text/images, tool definitions/arguments/grammars/results, and human responses so terminal outcomes remain durable; it does not send that content to an AI provider. A schema or grammar does not grant execution permission: the caller remains responsible for approving and safely executing its own tools. Provider configuration must not weaken a coding harness's sandbox or approval policy.