Published
Updated
PolicyLM-1.7B scores each policy category from 0 to 1
PolicyLM-1.7B prints custom-policy accuracy 0.842 at a 0.5 cutoff, and it returns a score from 0 to 1 for each category, up to 16 in one pass. The weights are Apache-2.0. The table has no Jev column.
Musubi Labs published PolicyLM-1.7B as an open-weight moderation classifier. A caller sends a message and a policy. The model returns a score from 0 to 1 for each category, up to 16 categories in one pass. Explicit mode uses categories the caller writes. Aegis mode uses NVIDIA Aegis 2.0, 23 categories, plus one score named “Prompt harmful.” The license is Apache-2.0. The checkpoint is 1.7 billion parameters, a fine-tune of BidirLM-1.7B-Embedding from Qwen3-1.7B-Base. The quickstart names revision v1.2.
The public post that links the card is @Secretof_forest, October 7, 2026, 03:49 UTC. We did not find an official Musubi Labs post. The figures on this page are the card.
A short message, with up to six categories, has a median of 35 ms on one 24 GB L4 and 22 ms on an H100. The card also names a laptop CPU and Apple silicon. The batched table is L4 at 35 ms and 39.3 messages a second, L40S at 34 ms and 127.2, and H100 PCIe at 22 ms and 134.1.
Custom policy
The custom-policy benchmark says every set except OR-Bench was also used while the model was built. Accuracy at a 0.5 cutoff:
| Model | Accuracy | Policy edits followed | H100 median |
|---|---|---|---|
| PolicyLM-1.7B | 0.842 | 0.528 | 22 ms |
| gpt-oss-safeguard-20B | 0.909 | 0.716 | 349 ms |
The same table prints accuracy 0.829 for CoPE-B-A4B and 0.722 for Granite Guardian 4.1-8B, and it does not print the other two columns for those rows. It has no Jev column.
Cutoffs on the card are explicit precision 0.335, marked as the default, and explicit balanced 0.275. Aegis precision is 0.69 and aegis balanced is 0.45. The policy and the message share 2,048 tokens, and the policy can use up to 1,662 of them. The card says child-safety enforcement is out of scope.
The card also prints AUROC for PolicyLM: WildGuardTest 0.947, XSTest 0.968, OR-Bench 0.867, PolyGuard 0.906, and RabakBench 0.835. Each of those columns has a higher score on another row. The figure this desk catalogs is the custom-policy 0.842. We did not rerun the table.
This site's reading
Editorial notes evaluating claims against primary sources, contextualizing findings alongside related implementations, and defining technical terms.
Verify
The public post that links the card is @Secretof_forest, October 7, 2026, 03:49 UTC, status 2107679726617923648. We did not find an official @musubilabs post. The numbers below are the card huggingface.co/musubilabs/policylm-1.7b, opened the same day.
The card describes an open-weight moderation classifier. The caller sends a message and a policy. The model returns a score from 0 to 1 for each category, up to 16 categories in one pass. Explicit mode uses categories the caller writes. Aegis mode uses NVIDIA Aegis 2.0, 23 categories, plus one score named Prompt harmful. The license is Apache-2.0. The size is 1.7B, a fine-tune of BidirLM-1.7B-Embedding from Qwen3-1.7B-Base. The quickstart names revision v1.2.
A short message, up to six categories, has a median of 35 ms on one 24 GB L4, and 22 ms on an H100. The card also names a laptop CPU and Apple silicon. The batched table is L4 at 35 ms and 39.3 messages a second, L40S at 34 ms and 127.2, and H100 PCIe at 22 ms and 134.1.
The custom-policy benchmark says every set except OR-Bench was also used in development. Accuracy at a 0.5 cutoff is PolicyLM 0.842, policy edits followed 0.528, H100 median 22 ms. gpt-oss-safeguard-20B is 0.909, 0.716, and 349 ms. CoPE-B-A4B is 0.829. Granite Guardian 4.1-8B is 0.722. The table has no Jev column.
Cutoffs on the card: explicit precision 0.335, the default, and balanced 0.275. Aegis precision 0.69 and balanced 0.45. The policy and the message share 2,048 tokens, with up to 1,662 tokens of policy. The card says child-safety enforcement is out of scope.
The AUROC cells for PolicyLM are WildGuardTest 0.947, XSTest 0.968, OR-Bench 0.867, PolyGuard 0.906, and RabakBench 0.835. Each of those columns has a higher score on another row. We did not rerun the table.
Compare
The catalog figure is the custom-policy accuracy, 0.842 at a 0.5 cutoff. The card says every set in that benchmark except OR-Bench was also used while the model was built. The table has no Jev column.
This call returns a score from 0 to 1 for each category in a policy. The card does not serve Choice, Noul, or POST /v1/systemone. Jev's Noul is a yes/no probability on one question. The two requests ask for different returns.
Terms
- PolicyLM-1.7B
- Musubi Labs' Apache-2.0 moderation classifier, a 1.7B fine-tune of BidirLM-1.7B-Embedding. A message plus a policy returns a score from 0 to 1 for each category, up to 16 in one pass.
- 0.842
- Custom-policy accuracy at a 0.5 cutoff on the PolicyLM card. Policy edits followed is 0.528. The H100 median on that row is 22 ms. The table has no Jev column. The card says every set except OR-Bench was also used in development.