Pokee-Isaac 28B Focuses on Ultra-Long Context

Pokee-Isaac 28B combines a 10-million-token context window, single-GPU deployment, and agentic tools, with API input priced at $0.15 per million tokens.

Pokee-Isaac 28B is a 28-billion-parameter agentic model capable of reasoning, planning, and using tools across a context window of up to 10 million tokens. Its proprietary architecture is not exclusively decoder-based, although the report does not explain its design in detail. Some of its weights were fine-tuned from Qwen3.6-27B.

The model can be installed in a VPC, on internal infrastructure, or on certain devices, with support announced for vLLM, SGLang, and OpenAI-compatible interfaces. Pokee states that a single RTX 4090 or 5090 is sufficient for private deployment, while measurements at the maximum context length were conducted on an NVIDIA B200.

According to the company, Pokee-Isaac scores 93.3% on RULER with a 10-million-token context window. On a B200, processing a prompt of that size reaches 137,200 tokens per second, with a 72.9-second time to first token and output generation holding at approximately 337 tokens per second.

In Pokee’s agentic evaluations, the model scores 70.94 on BFCL v4, slightly ahead of GPT-5.6 Luna at 70.61. It also ranks first in the comparison group on τ³-bench with a score of 0.662, but remains behind GPT-5.6 Luna on Terminal-Bench 2.1 and behind Luna and Gemini 3.5 Flash Lite on MCP-Atlas.

API access is priced at $0.15 per million input tokens and $1 per million output tokens. The model is currently text-only, with no support for image, audio, or video inputs.