Perplexity's pplx-decider-v1-27b scores 85.71% on a 7,210-row panel

Opened October 8, 2026, Decision Index 0.3 still ranks Perplexity Decider v1.1 first at 62.8, with Jev 1.13.0 as a reference at 60.1. The header says 115 models, updated 2026-10-07. Opened October 8, 2026, the README frontmatter lists license apache-2.0, and the model API lists private false and gated false. The usage section still says the repository is private. Decision Index 0.3, opened October 7, 2026, ranked Perplexity Decider v1.1 first at 62.8, with Jev 1.13.0 as a reference at 60.1. The quickstart opened the same day prices the Decisions API at $0.02 per million input tokens, with output free. Perplexity posted pplx-decider-v1-27b on October 1, 2026. On a fixed 7,210-row panel the chart prints 85.71% against jev-1.13.0 at 84.51% and Qwen3.8-27B at 74.76%. The October 1 quickstart priced that call at $0.04 per million input tokens. A Fastino chart the same week prints 56.39 for this model on Decision Index 0.2.1, which is a different suite. On October 4 Aravind Srinivas posted a FireRed Elite Four run: 592 ms median, 137 calls, estimated cost $0.028.

Perplexity posted pplx-decider-v1-27b on October 1, 2026. The chart on that post is a fixed panel of 7,210 rows, dated September 2026. Overall accuracy is 85.71% for pplx-decider-v1-27b, 84.51% for TypeSafe jev-1.13.0, and 74.76% for the open-weight Qwen3.8-27B.

The Hugging Face card says the weights are Apache-2.0, a fine-tune of Qwen/Qwen3.8-27B, about 49 GiB, with a 250k context. We did not load them.

What the eleven rows say

BenchmarkRowsJevQwen3.8-27Bpplx-decider
WinoGrande100090.70%73.10%83.30%
FinancialPhraseBank99976.98%75.68%84.18%
RAGTruth150077.27%61.53%88.80%
JudgeBench35078.57%68.86%78.29%
BBH75094.27%72.80%82.80%
JevBench public hard10173.27%72.28%70.30%
TabFact50089.80%78.60%90.60%
Circa50084.60%87.00%89.20%
Belebele50095.00%93.20%94.00%
TruthfulQA binary50092.00%82.80%85.40%
ContractNLI51077.45%80.78%80.78%
Overall721084.51%74.76%85.71%

Jev is ahead on WinoGrande, JudgeBench, BBH, the public-hard slice, Belebele, and TruthfulQA. pplx-decider is ahead on FinancialPhraseBank, RAGTruth, TabFact, and Circa. On ContractNLI the open base and the fine-tune print the same 80.78%, and Jev is at 77.45%. JudgeBench is close: 78.57% against 78.29%.

That public-hard row is 101 items. The Open-Jev hard tier is 111 tasks on another slice of public JevBench. The 85.71% overall is also not the September 28 Decision Index cell. That file lists Jev at 57.91 and AutoJev-27B at 56.40. This card does not say the weights are the checkpoint behind that AutoJev row.

What the hosted call costs

Opened October 1, the quickstart said the route is POST https://api.perplexity.ai/v1/decisions with model pplx-decider-v1-27b. Input on that reading was $0.04 per million tokens. Output was free. There is no per-request fee. The header it accepts is a bearer token. An x-api-key header is a 401. The org limit is 10 requests a second.

A request can hold 1 to 128 questions. A choice stops at 255 options. A score stops at 10 levels. Input has to stay under 262,144 tokens, and the body under 32 MiB. Images go in as base64 data URLs, cut into 32 by 32 tiles. The page says a request past 2,048 tiles, or about 60 seconds, returns 504.

The same page prints example latencies from a test on September 30, 2026. A few hundred tokens finished in under 2 seconds. About 90k tokens took 5 seconds, about 190k took 14, and a prompt near the cap took 23. The JSON sample next to those lines is an example payload. It is not a score on the panel.

A 4-bit timing, and a second index

On October 3 David Hendrickson wrote that a ramgpt EXL3 build of this model fits in 15.35 GB at 4.00 bits per weight. The medians he prints are 135 ms, 278 ms, and 430 ms at 128, 512, and 1,024 tokens. He says the card he is quoting has no accuracy test of that build against BF16, and that the official BF16 weights want about 49 GiB. He also restates the panel: 85.71%, 84.51%, 74.76%, and Jev still ahead on 6 of the 11. The download link was in the image alt text. We did not open it.

The same morning, a Fastino chart prints pplx-decider at 56.39 on Decision Index 0.2.1 and says Fastino ran that suite. The caption puts Jev at the public-board 57.91. That 56.39 is not a cell on the 7,210-row panel.

On October 4, 2026, at 10:10 UTC, Aravind Srinivas posted a Pokémon FireRed run. The post says the Elite Four and the Champion were cleared in one shot, with decision-making from Perplexity’s Decisions API. The stats in the post are a 592 ms median API response, a 987 ms p95, 96.4% of responses under one second, an estimated inference cost of $0.028, and 137 live API calls. The post does not say whether the median was measured on the client or on the server. It prints no win rate against another model, and it prints no Jev cell. The attachment is a video. We did not rerun the game.

v1.1, October 6

Perplexity posted pplx-decider-v1.1-27b on October 6, 2026, at 17:13 UTC. The post says the model is highest on Decision Index 0.3 and costs half as much as v1, at $0.02 per million input tokens. A follow-up the same minute says the context is 250k, the inputs are text and image, and links the weights and the docs. Aravind Srinivas wrote that Perplexity wins on the Hugging Face Decision Index and that the weights are open.

Opened October 7, the quickstart prices the Decisions API at $0.02 per million input tokens. Output is free. There is no per-request fee. The page names both pplx-decider-v1.1-27b and pplx-decider-v1-27b. The route is still POST https://api.perplexity.ai/v1/decisions. The page prices the API. It does not split the rate by model. The caps from the October 1 reading are still on the page: 10 requests a second, 1 to 128 questions, 255 choice options, 10 score levels, input under 262,144 tokens, a body under 32 MiB, images as base64, and 2,048 tiles of 32 by 32. The JSON sample on the page is an illustration. We did not send it.

The v1.1 card says the checkpoint updates v1 on the same backbone. It says a Decision Index score moves from 56.4 to 61.56, and that 61.56 is more than 3.5 points above Jev. The table on the card, in the order Jev, v1, v1.1:

AreaJevv1v1.1
Knowledge51.440.948.18
Language62.063.569.45
Retrieval55.454.961.26
Tools75.179.378.88
Arts37.739.444.66
Overall57.956.461.56

The card says the overall is a suite weighting. The five v1.1 area scores average to 60.486. The printed overall is 61.56. It attributes the gain to lifting the causal mask and to tasksource data. The native checkpoint is a Qwen3_5Model plus readout.safetensors, BF16, shape [255, 5120]. The base is Qwen/Qwen3.8-27B. The card says about 49 GiB. A model-size line says 26B parameters. The model name says 27B. The Decision Index 0.3 row labels the same name 28B. The card says full-attention layers use a saved noncausal mode. The card we read has no license line. It says the repository is private and that access is authenticated. The October 6 post says the weights are open. A request to the Hugging Face model API from this host returned a rate limit, so this page does not add a second check of that gate. We did not load the weights.

Opened October 8, 2026, the same README frontmatter lists license: apache-2.0. The usage section still says authenticated Hugging Face access to this private repository. The model API lists private false and gated false. The safetensors total is 26,085,330,160 parameters. lastModified on that response is 2026-10-05. The score table is still 61.56 against Jev at 57.9. We did not load the weights.

Opened October 8, 2026, the Decision Index space still says edition v0.3. The header says 115 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-07. The lede says 114 open reproductions. The head-to-head panel still lists this model first at 62.8, with Jev as a reference at 60.1. The full-results table prints =2 on the GLiDE row, the Jev row, and the Torchcast row. The v1.1 row still prints 62.8, then 45.8, 62.3, 64.7, 80.6, and 51.4, then 28B, 100.0%, and 104 ms. We did not rerun the index.

The Decision Index space, opened October 7, 2026, headers edition v0.3, 112 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-06. The full-score summary lists Jev at 60.1 as a reference, Perplexity Decider v1.1 (27B) at 62.8 in rank 1, Fastino GLiDE no-thinking (28B) at 60.2 in rank 2, and Torchcast Decision 27B at 59.9 in rank 3. The GLiDE row’s button says open weights are coming soon. The model id on that row is decision-27b-ckpt-nothink. The v1.1 ranking row prints full score 62.8, then 45.8, 62.3, 64.7, 80.6, and 51.4. Those five match the area tabs in the order Knowledge, Language, Retrieval, Tools, Arts. The same row prints an empty cell, 28B, 100.0%, and 104 ms. We did not keep the column titles for the empty cell, the percent, and the milliseconds. The row labels the model a full fine-tune of Qwen/Qwen3.8-27B, autoregressive.

On the Language tab, Jev’s reference is 58.3 and v1.1 is rank 1 at 62.3. On Retrieval, Jev’s reference is 61.4 and v1.1 is rank 1 at 64.7. The Knowledge list we read starts with Quyet-1.0-Large at 47.8, then v1.1 at 45.8. Tools starts with GLiDE no-thinking at 82.7, then v1.1 at 80.6. Arts starts with GLiDE at 51.7, then v1.1 at 51.4. We did not copy a Jev reference for Knowledge, Tools, or Arts. A search of the space for “Decider v1 (27B)” and for “Drex” returned no row. The formula on the page is 0.20 times public, plus 0.50 times same-skill private, plus 0.30 times new domains. The public part is 100 times the weighted mean of the area scores, with GSM8K rebuilt and ForecastBench retired.

The card’s 61.56 against 57.9 is the table above. The space’s 62.8 against the reference 60.1 is the edition 0.3 full score. The October 1 panel’s 85.71% is a third measurement. We did not call Perplexity, and we did not rerun these tables.

This site's reading

Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.

Verify

Perplexity, @perplexitydevs, October 1, 2026, 18:24 UTC, status 2105725611599954234. The attached chart is the image saved for this page. The subtitle says accuracy on a fixed 7,210-row panel, higher is better, September 2026.

The columns are Jev, named as TypeSafe jev-1.13.0, Qwen3.8-27B open weights, and pplx-decider-v1-27b. The overall row is 84.51%, 74.76%, and 85.71%. Row counts on the chart add to 7,210.

The Hugging Face card perplexity-ai/pplx-decider-v1-27b says Apache-2.0, base Qwen/Qwen3.8-27B, about 49 GiB, and a 250k context. We did not load the weights.

Opened October 1, the quickstart at docs.perplexity.ai said POST https://api.perplexity.ai/v1/decisions, model pplx-decider-v1-27b, $0.04 per million input tokens, output free, and no per-request fee. Auth is a bearer token. The page says an x-api-key header returns 401. The org cap is 10 requests a second.

The same page caps a request at 128 questions, a choice at 255 options, a score at 10 levels, and the input at under 262,144 tokens. The body cap is 32 MiB. Images are base64 data URLs in 32 by 32 tiles. Over 2,048 tiles, or about 60 seconds, the page says the call returns 504.

Timeout examples on that page, tested 2026-09-30, are under 2 seconds for a few hundred tokens, 5 seconds near 90k, 14 seconds near 190k, and 23 seconds near the cap. The sample JSON there is an example response.

David Hendrickson, @TeksEdge, October 3, 2026, 06:00 UTC, status 2106262976466489598, reports a ramgpt EXL3 build at 4.00 bits per weight and 15.35 GB, with medians 135 ms, 278 ms, and 430 ms at 128, 512, and 1,024 tokens. He says that card has no BF16-versus-EXL3 accuracy test, and that official BF16 wants about 49 GiB. The link was in the image alt text. We did not open it.

We did not call the API and we did not rerun the panel. TypeSafe's Master Customer Agreement section 2.3(f) forbids customers from publishing benchmarks of the Services. The Jev cells above are the ones on the chart.

Aravind Srinivas, October 4, 2026, 10:10 UTC, status 2106688463542333795. The post says a FireRed Elite Four and Champion were cleared in one shot on Perplexity's Decisions API. Median 592 ms, p95 987 ms, 96.4% of responses under one second, estimated cost $0.028, 137 live calls. The post does not say whether the median is client time or server time. The attachment is a video. We did not rerun the game.

Perplexity, @perplexitydevs, October 6, 2026, 17:13 UTC, status 2107519531711418597. The post says pplx-decider-v1.1-27b is available, highest on the new Hugging Face Decision Index 0.3, at half the v1 price, $0.02 per million input tokens. The follow-up, status 2107519543996494241, says 250k context, text and image, and links the weights and the docs. Aravind Srinivas, status 2107522748167954442, writes that Perplexity wins on the Hugging Face Decision Index and that the weights are open.

The v1.1 card, opened October 7, says the checkpoint updates v1 on the same backbone and moves a Decision Index score from 56.4 to 61.56, more than 3.5 points above Jev. The table is Knowledge 51.4, 40.9, 48.18; Language 62.0, 63.5, 69.45; Retrieval 55.4, 54.9, 61.26; Tools 75.1, 79.3, 78.88; Arts 37.7, 39.4, 44.66; overall 57.9, 56.4, 61.56, in the order Jev, v1, v1.1. The card says the overall is a suite weighting. It attributes the gain to lifting the causal mask and to tasksource data. The native checkpoint is a Qwen3_5Model plus readout.safetensors, BF16, shape [255, 5120]. The base is Qwen/Qwen3.8-27B. The card says about 49 GiB, a model-size line of 26B parameters, and full-attention layers in a saved noncausal mode. The card we read has no license line. It says the repository is private and that access is authenticated. A request to the Hugging Face model API from this host returned a rate limit, so the gate was not checked past that sentence. We did not load the weights.

Opened October 8, 2026, the README frontmatter lists license apache-2.0. The usage section still says authenticated Hugging Face access to this private repository. The model API lists private false and gated false. The safetensors total is 26,085,330,160 parameters. lastModified on that response is 2026-10-05. The score table is still 61.56 against Jev at 57.9.

The quickstart, opened October 7, prices the Decisions API at $0.02 per million input tokens, output free, no per-request fee, and names both pplx-decider-v1.1-27b and pplx-decider-v1-27b. The page prices the API. It does not split the rate by model. The caps from the October 1 reading are still on the page. The JSON sample is the page's illustration.

The Decision Index space, opened October 7, headers v0.3, 112 models, Jev jev-1.13.0, reproductions on one NVIDIA RTX PRO 6000, updated 2026-10-06. The full-score summary lists Jev at 60.1 as a reference, Perplexity Decider v1.1 (27B) at 62.8 in rank 1, Fastino GLiDE no-thinking (28B) at 60.2 in rank 2, and Torchcast Decision 27B at 59.9 in rank 3. The GLiDE button says open weights are coming soon. The v1.1 row prints 62.8, then 45.8, 62.3, 64.7, 80.6, and 51.4, matching the area tabs Knowledge, Language, Retrieval, Tools, Arts. The same row prints an empty cell, 28B, 100.0%, and 104 ms. We did not keep the column titles for those three cells. Language lists Jev at 58.3 and v1.1 at 62.3, rank 1. Retrieval lists Jev at 61.4 and v1.1 at 64.7, rank 1. The Knowledge list we read starts at Quyet-1.0-Large 47.8, then v1.1 at 45.8. Tools starts at GLiDE no-thinking 82.7, then v1.1 at 80.6. Arts starts at GLiDE 51.7, then v1.1 at 51.4. We did not copy a Jev reference for Knowledge, Tools, or Arts. A search for Decider v1 (27B) and for Drex returned no row. The formula on the page is 0.20 times public, plus 0.50 times same-skill private, plus 0.30 times new domains. The public part is 100 times the weighted mean of the area scores, with GSM8K rebuilt and ForecastBench retired. We did not call the API.

Compare

Jev is the higher cell on six of the eleven named sets: WinoGrande, JudgeBench, BBH, JevBench public hard, Belebele, and TruthfulQA. pplx-decider is higher on FinancialPhraseBank, RAGTruth, TabFact, and Circa. ContractNLI is 80.78% for both Qwen3.8-27B and pplx-decider, against Jev at 77.45%.

The JevBench public hard row here is 101 items, Jev at 73.27%. Open-Jev's hard tier is 111 tasks on a different public slice. Decision Index 0.2.1 lists Jev at 57.91 on the September 28 board. AutoJev-27B at 56.40 there is a different composite. This note does not treat these weights as the same bytes as denis-pplx/autojev-27b.

Fastino's October 3 chart prints pplx-decider at 56.39 and says Fastino ran the full Decision Index suite. That chart is on the GLiDE page. Hendrickson's medians are a local EXL3 report. They are not the hosted timeouts in the quickstart, and they are not an accuracy comparison against BF16.

The October 4 FireRed post prints 137 calls and $0.028 for one cleared run. It prints no Jev cell and no second model. The 592 ms median is a game loop, not a cell on the 7,210-row panel.

62.8 is the space's Decision Index 0.3 full score, with Jev 1.13.0 as a reference at 60.1. 61.56 is the v1.1 card's table against Jev at 57.9. 85.71% is the October 1 panel. 56.39 is Fastino's October 3 chart on edition 0.2.1. The ranking row labels the model 28B. A line on the card says 26B parameters. The model name says 27B.

Terms

pplx-decider-v1-27b
Perplexity's decision model, posted October 1, 2026. The card says it is an Apache-2.0 fine-tune of Qwen3.8-27B with a 250k context. The October 1 quickstart priced the hosted route at $0.04 per million input tokens.
pplx-decider-v1.1-27b
Perplexity's October 6 update of the same backbone. The October 7 quickstart prices the Decisions API at $0.02 per million input tokens and names this model. Decision Index 0.3 ranks it first at 62.8. The card's own table prints 61.56 against Jev at 57.9.
7,210-row panel
The chart's fixed set, dated September 2026. Eleven named benchmarks plus an overall row. Jev 84.51%, Qwen3.8-27B 74.76%, pplx-decider-v1-27b 85.71%.

Sources

  1. Perplexity, October 1
  2. pplx-decider-v1-27b
  3. Perplexity Decisions quickstart
  4. David Hendrickson, October 3
  5. Sahibzada Allahyar, October 3 chart
  6. Aravind Srinivas, FireRed, October 4
  7. Perplexity, v1.1, October 6
  8. Perplexity, v1.1 weights, October 6
  9. Aravind Srinivas, Decision Index, October 6
  10. pplx-decider-v1.1-27b
  11. Decision Index space, opened October 7