Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It is intended for harder reasoning and evaluation workloads, with automatic prompt caching and a one-million-token context window.
Features
Context
1.0M
Max Output
1000.0K
Date Added
Aug 31, 2026
Input:
$5.00/1M
Output:
$5.00/1M
Cache:
Read $0.50/1M
Est./msg:
$0.0075
Not included in subscription
View Providers
GLM-5.3 is Z.AI's open-weight reasoning model for complex software engineering, autonomous agents, vulnerability research, and long-horizon tasks. This text-only deployment runs inside a Phala Trusted Execution Environment with Redpill attestation and signed completion receipts.
Features
Context
1.0M
Max Output
131.1K
Parameters
744B / 40B
Date Added
Aug 31, 2026
Input:
$1.40/1M
Output:
$4.40/1M
Cache:
Read $0.26/1M
Est./msg:
$0.0036
Not included in subscription
View Providers
IBM Granite 4.2 8B is an Apache 2.0-licensed dense model with native step-by-step reasoning and specialized training for agentic work. It can plan before acting, sequence tools, navigate codebases, work in terminals, and verify results across coding, search, mathematics, science, and complex instruction-following tasks.
GLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters, served by Phala inside a Trusted Execution Environment with Redpill attestation and signed completion receipts.
Hy4 Preview is Tencent's 770B-parameter mixture-of-experts model with 49B active parameters. It is designed for coding agents, complex tool-use workflows, and productivity tasks, with a 1M-token context window and configurable reasoning effort.
Abliteration.ai's unrestricted multimodal reasoning model is weight-modified to reduce refusals and answer prompts the original model might reject. It supports text and image input, structured output, automatic prompt caching, and a 262K-token context window.
Abliteration.ai's unrestricted large text reasoning model is derived from GLM-5.2 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
This is the same underlying model as Qwen3.8 Max, exposed under its architecture-based 2.4T A95B name for easier discovery. It uses the identical routing, pricing, capabilities, and non-thinking mode.
Qwen3.8 Flash is Alibaba's latest fast multimodal model, with a million-token context window for coding, agentic workflows, visual understanding, long documents, codebases, and videos.
Qwen3.8 27B is an open-weight dense vision-language model from Alibaba for reasoning, coding, professional workflows, multimodal interaction, tool use, and structured output. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
An experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model. This route is served directly by DeepSeek, so privacy and logging guarantees are limited.
Qwen3.5 0.8B is a lightweight open-weight multimodal model from Alibaba for fast reasoning, visual understanding, tool use, and JSON output.
Qwen3.5 2B is a small open-weight multimodal model from Alibaba for efficient reasoning, coding, visual understanding, tool use, and JSON output.
Qwen3.5 4B is a compact open-weight multimodal model from Alibaba for reasoning, coding, visual understanding, tool use, and structured output.
Qwen3.8 27B is an open-weight multimodal model from Alibaba for coding, visual understanding, tool use, and structured output. This variant keeps thinking disabled for faster direct responses.
Qwen3.8 27B is an open-weight multimodal model from Alibaba for reasoning, coding, visual understanding, tool use, and structured output. This variant enables thinking by default.
Dots Studio's open-weight multimodal Mixture-of-Experts model activates 16B of 280B parameters for long-context reasoning, coding, visual and document understanding, tool use, and long-horizon agent workflows. Prompts and completions of this model may be logged.