What this is
A study from the laugh.so team tracked three open-source models—Tulu 3 (based on Meta's Llama 3.1 70B), OLMo 3.1 32B, and Alibaba's Qwen2.5—across 11 post-training stages (the secondary tuning applied after pre-training to teach a model to understand and respond to human instructions). Using 100 joke-telling prompts, 64 human raters, and 2,330 pairwise comparisons, the team reached a clear conclusion:
Post-training does make models funnier—5 of 7 training steps were judged funnier, jokes shortened by an average of 10–20 words, and punchlines arrived faster. But the cost is a collapse in diversity: across 6 steps, the models' outputs to the same prompt grew increasingly similar—ask for 8 different jokes, you get 8 variants of the same one. The drop from Qwen2.5's base version to its instruction-tuning stage was the single largest diversity loss.
The researchers call this trade-off the "humor tax."
Industry view
The methodology deserves credit—this is one of the few studies that breaks down post-training's effects stage by stage, and it covers open-source models our readers actually use.
But there are cooler voices. Some researchers note that "humor" is just the tip of the iceberg when it comes to post-training's effects; the more important question is whether this "diversity decline" also hits more practical tasks like reasoning, writing, and code. One product manager told us privately: this perfectly explains why AI assistants on the market all sound the same—everyone is using similar RLHF (Reinforcement Learning from Human Feedback) methods for post-training, optimizing models to be "safe, polite, plausible-sounding"—but creativity is being systematically sanded down.
The study also tested a counter-intuitive fix: have the model "plan a line or two" before telling the joke. Diversity dropped across all 4 models, with no stable improvement in humor—"drafting" doesn't make AI funnier. "Comedy persona" prompts recovered some diversity but only made 2 models funnier.
Impact on regular people
For enterprise IT: If you're using AI for marketing copy, brand slogans, or social content, be aware: the more "finished" a model is, the more homogenized its output. If you need diversity, you may need to deliberately retain the base model or an earlier training stage.
For individual professionals: Why does brainstorming with AI always give you variations of the same idea? It's not just a prompting problem—the diversity was already pruned during post-training. Switching to an earlier-stage model or varying the random seed (the randomness injected at generation time) several times may work better than tweaking prompts.
For the consumer market: All AI assistants converging into the same product is no coincidence. When the industry relies on similar post-training pipelines, "sounding like AI" becomes a new homogenization—future differentiation will likely come from product form and data access, not the model itself.