Why LLMs can't write well and how Sherpa fixes it

@RohanNayak2
Rohan Nayak@RohanNayak2
45 views Oct 01, 2026 ~9 min read
Advertisement

As we pivoted to AI-written shows, countless people told us this would kill Pocket FM. After six long months, I almost believed them. But then our ecosystem exploded, and not despite AI writing but because of it. Within <2 years, our AI has helped writers create 161 blockbusters on Pocket, including a $100M IP.

Media image

LLMs being terrible at writing forced us to build our own custom models and harness. Starting from scratch was demoralizing at times. It took an enormous amount of trial and error, and a golden dataset produced by 550k writers and 100M+ users on the platform.

I'll go over how we went from nothing to a beautifully calibrated writing model, why Pocket FM is the perfect testing ground, and the secret sauce behind Sherpa's ability to write blockbusters.
_______________________________________________________________________
Ever since ChatGPT was released, everyone has pointed out that it sounds too AI. When you ask an LLM to write anything, whatever its length, em-dashes appear, short, punchy sentences make their way into the writing, and the output reads like one big, boring piece of work.

LLMs were fundamentally designed for logical reasoning, answering math queries, helping with coding, and retrieving encyclopedic knowledge. Nothing in their training teaches them to write something a reader would willingly read or listen to. This is not the models' fault.

Reinforcement learning works best when output is tied to a clear success signal, training the model if something worked. In coding, the code either works and does what you asked for, or it doesn't. The answer is certain.

When you write something, you can't tell if it's any good. You don't know if the reader dropped off after the first sentence or if they finished it. Even worse than that, ask 10 people about a specific book or article they've read and you'll get 10 different answers.

Owning creation and distribution puts Pocket FM in a very unique position. We monitor usage minute by minute. We know when people stop listening, whether they come back for the next episode, and if they stay for the long run. There are no better signals than that. Whatever leads to a user dropping off, Sherpa learns to avoid; and whatever keeps the user engaged with a specific show, Sherpa leverages for new shows.

Media image

The average user on Pocket FM listens for 150+ minutes a day, almost 3x that on YouTube and ~50% more than on Netflix. I think we're at the very beginning of a new wave of extremely entertaining, immersive and hyper personalized content, which will only get better as we get more data.

Why AI writing sounds like AI writing

Dostoevsky was an incredible writer because he showed contradictions and emotions in his characters' behavior and rants. Oscar Wilde and Fitzgerald because they embellished their witty writings so beautifully you can't help but keep turning the pages. Conan Doyle because of his ability to conceal information from the reader until the very end.

General-purpose LLMs were built to do the exact opposite of this, with no way to become better at it. They give answers directly, write in neutral language, and overuse resources to enhance effectiveness of delivery. Similarly, good AI assistants are trained to lead you to the right conclusion quickly. All these fundamental pieces of wiring permeate their writing, and that's why they are bad at it.

Creative writing requires tension, emotion, creation and slow resolution of conflict, character traits that go beyond one-liners, and variations in pace. Users must be led to wrong conclusions while receiving clues that keep them hooked.

Even a reasoning model that spends longer thinking is generally working toward resolving the problem, rather than being rewarded for the audience’s experience of getting there. Solving a mystery and writing a mystery are two very different things.

We failed by trying to tweak what existed

At the beginning, we thought we could get away with using what was available. We tried off-the-shelf models and tried refining them in the hope of repairing their soulless storytelling.

We experimented with prompting increasingly strong models directly, maintaining rolling summaries as the story evolved, and using retrieval and knowledge graphs to answer queries.

By prompting directly, you reach episode 100 with 10-20 inconsistencies. Rolling summaries essentially send important details into the ether, impairing the story thereafter. A larger context window helps, but models can still struggle to use information accurately in long contexts.

Aside from the actual output, queries are the best measure of a model's understanding of what it has written. Early on, we realized that no matter how much effort we put into any model, it failed to answer multi-hop narratological queries, which require details from multiple episodes. This revealed the limitations with the state of the technology.

We concluded this path was unsustainable and decided to build our own technology... Sherpa, rightfully named to guide a writer through the hardest parts of storytelling

Technical reasons why Sherpa writes a story you want to listen to

Sherpa is a long form, serialized creative writing model harness. It's made of three components, which I'll describe in more detail after this intro:

  • A planner decides what each arc and scene needs to accomplish, and about character development and relationships evolution.
  • A Narrative World Model keeps a structured record of what has happened, what each character knows, and which promises the story still owes the audience.
  • A prose engine drafts against that plan and state, while critics check character behavior, continuity, and scene quality. Writer edits and audience response then help improve the next iteration.
  • Results are very encouraging. When Sherpa and each competitor wrote a story from a blank page, blinded pairwise judgments preferred Sherpa in 65.6% to 97.1% of cases, depending on the competitor. Given Sherpa's reinforcement learning and Pocket's growing dataset, there's no reason why the gap shouldn't continue to widen.

    Media image

    Memory - Narrative World Model (NWM)

    Language models have no memory between calls, so everything they know about a story has to fit in the prompt. Longer context windows don't solve this because models make uneven use of information buried in the middle of a long prompt. Retrieval over raw text doesn't solve it either. A search can return the passage that says where an object was, not where it is now.

    When an episode is finalized with Sherpa, a model extracts structured records from it, each tied to the passage that supports it. The fields are built for fiction: what characters know and don't know yet, when an event happened versus when the audience learns it, which setups are waiting for a payoff. Every fact is valid from the chapter that established it until the chapter that changes it. If a writer edits an earlier chapter, everything that depended on it is rebuilt.

    To answer questions, NWM searches the graph by exact wording and meaning at the same time, pulls in the direct connections of each result, and hands the model a small, focused packet.

    Its episode-aware retrieval surfaces only relevant canon, enabling multi-hop reasoning without future-story leakage. In an internal 60-item benchmark, Sherpa NWM achieved 86.7% accuracy (52/60) at $0.7793 per run. This is the same accuracy as Claude Code's memory harness with opus 5.x class models, at approximately 21× lower cost ($16.2172 per run). It also outperformed prose-only retrieval by 20 percentage points (66.7% → 86.7%) while costing substantially less.

    Media image

    Prose Engine: so difficult to crack we had to build our own model

    Reinforcement learning needs a reward signal, and prose doesn't have an obvious one. A paragraph has no test that says whether it's Nobel-prize-quality or filler.

    Sherpa gives each part of the writing job to a different model. Each character is an agent, running on on our custom, post-trained models that proposes what it intends to do next based on its persona, the story's current goal, recent events, and the parts of the world relevant to it. A separate critic model reviews every proposal and sends back anything vague, implausible, out of character, repetitive, or not moving the story forward. A third model, the narrator, then chooses which proposals fit the scene and writes them into the next paragraph. Only the actions that reach the page are added to the story's memory.

    On top of this, we're training a proprietary prose model with reward functions built for serialized fiction. We penalize purple prose, repetition, and over-explanation, and reward natural dialogue, voices, and tension.

    Signals we measure prose on:
    - Does this contradict established events or character knowledge?
    - Is the action plausible, in character, and moving the scene forward?
    - Do writers prefer this dialogue, voice, and pacing to an alternative?
    - Do listeners finish the episode and return for later ones?

    These signals operate on different timescales. A sentence can sound good but break canon, and a cliffhanger can improve next-episode starts but weaken a whole arc. Sherpa needs to optimize across these signals rather than treating retention as a single score for “good writing" because of how subjective the latter is.

    Sherpa’s Prose Engine is already outperforming most of the open and commercial models we evaluated in our internal fiction benchmarks, with prose preference scores of 69% against Sonnet 5 and 94% against Gemini 3.1 Pro. Each bar shows preference for Sherpa’s writing over the named model, judged in both presentation orders, with 50% representing parity. These are early results from ongoing post-training, and we expect further gains as training continues.

    Media image

    Hierarchical Story Planner

    Story structure is a property of the whole show, and language models only generate the next piece. Nothing in next-word prediction knows that a clue planted in episode 5 owes a payoff in episode 60, or that this chapter needs a goal, a conflict, and an outcome. In a 100-page story from a frontier model, an AI editor sampling just five chapters flagged 52 structural problems: progression that didn't work, and chapters the story didn't need.

    Sherpa's planner works down from the overall narrative through arcs and episodes and gives every scene a job, including what it must leave unresolved, so episode 20 can prepare episode 80's reveal without giving it away. Every setup is stored in the story memory as an open promise until it pays off. A goal generator keeps the story moving: when a goal is met or stalls, it sets a new one that doesn't repeat any earlier goal. On the same 100-page test, structural problems fell from 52 to 31, and removing goal updates alone raised chapter-level critiques by nearly half.

    Media image

    _______________________________________________________________________
    Everything is in place for Sherpa to help writers create billion-dollar IPs in the next few years. Sherpa keeps getting better as we onboard more creators and users onto the platform.

    There's a lot of work ahead, and many issues remain unresolved; perfecting Sherpa is one of these, and so is creating the logistical conditions for writers to best use it.

    I'm extremely excited about what we are doing at Pocket FM. We are helping creators earn a living by telling stories that would have required a multimillion-dollar budget a few years ago.

    It's still early in what will be a $1T opportunity.

    Actions
    What You Can Do
    • Export as PDF or Markdown
    • Batch Export to Notion
    • Bookmark & Highlight
    • LinkedIn & Instagram Carousel Maker
    Create Free Account

    Includes 7-day Premium trial

    Advertisement