Every provider in this category restricts what its models will do. The restrictions overlap heavily at the core and diverge at the edges, and the divergence is where product plans get derailed.
The first is a design decision you can read about in advance and plan around. The second is an error, and every provider has some rate of it. Complaints online frequently conflate the two, which makes anecdotal reports a poor guide to how a provider will behave on your workload.
Over-refusal rates also move with each model release, in both directions, so a description written six months ago may no longer hold.
The core restrictions are broadly consistent: material that sexualises children, serious weapons uplift, targeted harassment, and helping with clearly criminal operations.
If your use case runs into these, the answer will be no everywhere, and provider shopping is not the solution.
Adult content, medical and legal specificity, security research, political persuasion, and content about real named people are all handled differently across providers and tiers.
These are the areas worth testing directly rather than assuming, because published policies rarely give enough granularity to predict behaviour.
Some providers permit uses under a commercial agreement that are restricted on consumer tiers, subject to review. Regulated industries sometimes have specific arrangements.
If a use case is central to your product and sits near a line, this is a sales conversation before it is a technical one.
Running weights yourself removes provider-side enforcement, but the licence still governs permitted use, and the law applies regardless of what a model will produce.
Open weights are not a route around legal constraints, only around one company's product policy.
Every major provider publishes one. Anthropic's is at anthropic.com/legal/aup, and the others maintain equivalents. They are shorter than most terms of service and worth reading in full if you are building something.
Three things to look for specifically: whether your industry or use case is named, what obligations fall on you when you deploy the model to your own users, and what happens if a customer of yours misuses it. That last point catches people out, since responsibility for downstream use typically sits with the developer rather than the provider.
Policies also change. A use permitted at launch may be restricted later, or the reverse. For anything a product depends on, this is worth rechecking periodically rather than once.
Rephrasing helps more often than it should, because over-refusals frequently trigger on wording rather than intent. Stating the professional context plainly, and being specific about what you want and why, resolves a substantial share of them.
Where it does not, providers have feedback routes, and reporting an incorrect refusal is genuinely useful to them. Where a legitimate use is consistently blocked, testing another provider is reasonable, since the divergence at the edges is real.
What does not work is trying to disguise a request that falls inside a documented restriction. That is a policy violation on most providers' terms regardless of whether it succeeds, and it can affect account standing.
A strictness ranking would need a standard set of borderline prompts, run across providers, scored consistently and repeated after every release. No such measurement is published, and informal comparisons are dominated by which prompts happened to be tried.
More importantly, strictness is not a virtue or a fault in itself. A provider that declines more is better for some deployments and worse for others, and the right question is whether a specific provider permits your specific use, which its policy answers directly.