{"object":"list","data":[{"id":"deepseek/deepseek-v4-flash-0731","object":"model","created":1785456000,"owned_by":"deepseek-ai","name":"DeepSeek V4 Flash","description":"DeepSeek V4 Flash is a 284B mixture-of-experts model activating roughly 13B parameters per token, with a 1M context window. Compressed latent attention makes a very long prompt cost far less memory than its length suggests.","canonical_slug":"deepseek/deepseek-v4-flash-0731","hugging_face_id":"deepseek-ai/DeepSeek-V4-Flash-0731","context_length":1048576,"max_completion_tokens":32768,"quantization":"mxfp4","pricing":{"prompt":"0.00000002","completion":"0.00000004","request":"0","image":"0","input_cache_read":"0.000000006","input":0.02,"output":0.04},"architecture":{"modality":"text->text","input_modalities":["text"],"output_modalities":["text"],"tokenizer":"deepseek","instruct_type":"deepseek"},"top_provider":{"context_length":1048576,"max_completion_tokens":32768,"is_moderated":false},"per_request_limits":null,"supported_parameters":["temperature","top_p","top_k","min_p","max_tokens","stop","seed","frequency_penalty","presence_penalty","repetition_penalty","tools","tool_choice","response_format","reasoning_effort","include_reasoning"],"default_parameters":{}},{"id":"qwen/qwen3.6-35b-a3b","object":"model","created":1785314439,"owned_by":"qwen","name":"Qwen3.6 35B A3B","description":"Qwen3.6 35B-A3B is a mixture-of-experts model with 35.9B total parameters and roughly 3B active per token, giving the quality of a large model at the speed of a much smaller one. It accepts text and image input, supports a 262K context window, and handles tool calling and structured output. Well suited to coding agents and long-context work where a whole repository or document set has to stay in the conversation.","canonical_slug":"qwen/qwen3.6-35b-a3b","hugging_face_id":"Qwen/Qwen3.6-35B-A3B","context_length":262144,"max_completion_tokens":32768,"quantization":"awq-int4","pricing":{"prompt":"0.0000001","completion":"0.0000009","request":"0","image":"0","input_cache_read":"0.000000025","input":0.1,"output":0.9},"architecture":{"modality":"text+image->text","input_modalities":["text","image"],"output_modalities":["text"],"tokenizer":"unknown","instruct_type":"qwen"},"top_provider":{"context_length":262144,"max_completion_tokens":32768,"is_moderated":false},"per_request_limits":null,"supported_parameters":["temperature","top_p","top_k","max_tokens","stop","stream","tools","tool_choice","seed","presence_penalty","frequency_penalty","response_format"],"default_parameters":{}},{"id":"qwen/qwen3-14b","object":"model","created":1745884800,"owned_by":"qwen","name":"Qwen3 14B","description":"Dense 14B instruct model from Alibaba with a hybrid thinking mode that can be switched per request. Competitive with much larger models on reasoning, maths and code, and the point in the Qwen3 range where quality stops being the constraint for most production work.","canonical_slug":"qwen/qwen3-14b","hugging_face_id":"Qwen/Qwen3-14B","context_length":131072,"max_completion_tokens":16384,"quantization":"awq-int4","pricing":{"prompt":"0.00000009","completion":"0.00000022","request":"0","image":"0","input_cache_read":"0.0000000225","input":0.09,"output":0.22},"architecture":{"modality":"text->text","input_modalities":["text"],"output_modalities":["text"],"tokenizer":"Qwen","instruct_type":"qwen"},"top_provider":{"context_length":131072,"max_completion_tokens":16384,"is_moderated":false},"per_request_limits":null,"supported_parameters":["temperature","top_p","top_k","min_p","max_tokens","stop","seed","frequency_penalty","presence_penalty","repetition_penalty","logprobs","top_logprobs","tools","tool_choice","response_format"],"default_parameters":{}},{"id":"qwen/qwen3-8b","object":"model","created":1745884800,"owned_by":"qwen","name":"Qwen3 8B","description":"Dense 8B instruct model from Alibaba. Strong general reasoning and multilingual performance for its size, with a hybrid thinking mode. A good default for agent loops and classification where an 8B is enough and latency matters.","canonical_slug":"qwen/qwen3-8b","hugging_face_id":"Qwen/Qwen3-8B","context_length":40960,"max_completion_tokens":8192,"quantization":"awq-int4","pricing":{"prompt":"0.00000005","completion":"0.00000015","request":"0","image":"0","input_cache_read":"0.0000000125","input":0.05,"output":0.15},"architecture":{"modality":"text->text","input_modalities":["text"],"output_modalities":["text"],"tokenizer":"Qwen","instruct_type":"qwen"},"top_provider":{"context_length":40960,"max_completion_tokens":8192,"is_moderated":false},"per_request_limits":null,"supported_parameters":["temperature","top_p","top_k","min_p","max_tokens","stop","seed","frequency_penalty","presence_penalty","repetition_penalty","logprobs","top_logprobs","tools","tool_choice","response_format"],"default_parameters":{}},{"id":"qwen/qwen3-8b:free","object":"model","created":1745884800,"owned_by":"qwen","name":"Qwen3 8B (free)","description":"Dense 8B instruct model from Alibaba. Strong general reasoning and multilingual performance for its size, with a hybrid thinking mode. A good default for agent loops and classification where an 8B is enough and latency matters.","canonical_slug":"qwen/qwen3-8b","hugging_face_id":"Qwen/Qwen3-8B","context_length":40960,"max_completion_tokens":8192,"quantization":"awq-int4-free","pricing":{"prompt":"0","completion":"0","request":"0","image":"0"},"architecture":{"modality":"text->text","input_modalities":["text"],"output_modalities":["text"],"tokenizer":"Qwen","instruct_type":"qwen"},"top_provider":{"context_length":40960,"max_completion_tokens":8192,"is_moderated":false},"per_request_limits":null,"supported_parameters":["temperature","top_p","top_k","min_p","max_tokens","stop","seed","frequency_penalty","presence_penalty","repetition_penalty","logprobs","top_logprobs","tools","tool_choice","response_format"],"default_parameters":{}}]}