Skip to content

New Text to Speech experiment - #888

Open
dkotter wants to merge 41 commits into
WordPress:developfrom
dkotter:feature/text-to-speech
Open

dkotter wants to merge 41 commits into
WordPress:developfrom
dkotter:feature/text-to-speech

Conversation

@dkotter

@dkotter dkotter commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

What?

Adds a new experiment, Text to Speech, that allows editors the ability to on-demand generate speech for a post, including the title and post content. This is then displayed as a player on the front-end or that can be turned off on a post by post basis.

Why?

Adding support for Text to Speech allows sites to provide their users with the ability to listen to content instead of having to read the content.

How?

  • Registers two new Abilities, ai/speech-generation and ai/speech-import. These can be used to generate speech from text and import that speech as an audio file into the Media Library
  • Add shared functionality across multiple classes that can take in some text content, chunk that content down (if needed, to stay under API limits. Note only works if the format is MP3) and send that content off to generate speech for. It then combines all the returned speech data into a single audio file which we import into the Media Library. This functionality is used both if someone directly uses the Abilities above and when audio generation is manually triggered in the admin
  • Add a REST endpoint that is used to track audio progress, allowing audio generation to happen via cron events. This allows longer running processes and processing of audio even if someone navigates away

Use of AI Tools

AI assistance: Yes
Tool(s): Claude Code
Model(s): Opus 4.8
Used for: Help with initial planning and implementation. Refinement, review and testing done by me

Testing Instructions

  1. Requires an AI Provider that supports Text to Speech. Easiest way to test for now is to download the OpenAI Provider plugin from this PR or Google Provider plugin from this PR
  2. Turn on the Text to Speech experiment
  3. Go to a post and find the Text to Speech sidebar panel. Click on the Generate Audio button
  4. Audio generation will take a bit but is run in the background via cron jobs. Easiest is to just stay on the page until it is done but you can navigate away and come back later
  5. Once finished, you should see an audio player render in that panel, along with a toggle to display the audio player on the front-end and a button to regenerate audio
  6. View the front-end and ensure the audio player shows there

Screenshots

Initial state of the Text to Speech panel in the editor sidebar State of the Text to Speech panel in the editor sidebar after audio has been generated Audio controls rendering on the front-end

Changelog Entry

Added - New experiment, Text to Speech, that allows you to generate audio for a specific post and then render audio controls on the frontend.

Open WordPress Playground Preview

dkotter added 17 commits July 16, 2026 15:04
…t to the AI Client to generate speech. Set up to be used for both Ability calls and background jobs
… Fires individual jobs, via cron, that will chunk content down and turn those into audio files, combining all files at the end
… a string of text or post content from a specific post ID
…nd imports it into the media library as an MP3 file
…he admin to make things look better. Add a better loading state and an icon to the button
@dkotter dkotter added this to the 1.3.0 milestone Jul 21, 2026
@dkotter dkotter self-assigned this Jul 21, 2026
@dkotter
dkotter requested a review from a team July 21, 2026 19:09
@dkotter dkotter added the [Status] Blocked Used to indicate unable to move forward label Jul 21, 2026
@dkotter
dkotter requested a review from jeffpaul as a code owner July 21, 2026 19:09
@github-actions

github-actions Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: dkotter <dkotter@git.wordpress.org>
Co-authored-by: saarnilauri <laurisaarni@git.wordpress.org>
Co-authored-by: jeffpaul <jeffpaul@git.wordpress.org>
Co-authored-by: drzraf <drzraf@git.wordpress.org>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@codecov

codecov Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 77.84145% with 232 lines in your changes missing coverage. Please review.
✅ Project coverage is 81.24%. Comparing base (a880957) to head (992a8c2).

Files with missing lines Patch % Lines
...udes/Experiments/Text_To_Speech/Text_To_Speech.php 69.37% 49 Missing ⚠️
...es/Experiments/Text_To_Speech/Speech_Generator.php 34.37% 42 Missing ⚠️
includes/Abilities/Speech/Generate_Speech.php 78.35% 29 Missing ⚠️
includes/Abilities/Speech/Import_Base64_Audio.php 84.44% 28 Missing ⚠️
...ncludes/Experiments/Text_To_Speech/Job_Manager.php 89.05% 22 Missing ⚠️
...udes/Experiments/Text_To_Speech/Voice_Resolver.php 76.34% 22 Missing ⚠️
...des/Experiments/Text_To_Speech/REST_Controller.php 85.57% 15 Missing ⚠️
includes/helpers.php 71.42% 10 Missing ⚠️
...udes/Experiments/Text_To_Speech/Audio_Combiner.php 77.14% 8 Missing ⚠️
...des/Experiments/Text_To_Speech/Content_Chunker.php 89.18% 4 Missing ⚠️
... and 1 more
Additional details and impacted files
@@              Coverage Diff              @@
##             develop     #888      +/-   ##
=============================================
- Coverage      81.53%   81.24%   -0.30%     
- Complexity      3071     3306     +235     
=============================================
  Files            129      138       +9     
  Lines          12253    13300    +1047     
=============================================
+ Hits            9991    10806     +815     
- Misses          2262     2494     +232     
Flag Coverage Δ
unit 81.24% <77.84%> (-0.30%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@drzraf

drzraf commented Aug 24, 2026

Copy link
Copy Markdown

I would also suggest one missing feature: the ability to easily filter by post-type (or just pass the post to allow for more granular control [sticky/date/...]). The vast majority of websites would likely want to selectively use it based on specific post type.

@dkotter

dkotter commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Feel free to merge, adapt, or ignore it.

Thanks @saarnilauri! Made a few changes but looked good so I've merged that in now

@dkotter

dkotter commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Most TTS (for accessibility etc...) distinguish between voice and language. If you choose a voice, you generally also have a distinct language setting alongside

@drzraf I don't believe OpenAI supports passing in a language (or even Google), they just infer the language from the text you send. ElevenLabs does allow you to pass the language code but it also will default to using the language of the text you send.

I guess I see no reason to complicate this further and introduce a setting to set the language when we can just rely on the LLM to match the language of the content given.

I would also suggest one missing feature: the ability to easily filter by post-type (or just pass the post to allow for more granular control [sticky/date/...]). The vast majority of websites would likely want to selectively use it based on specific post type.

I guess not sure what the request is here? Right now you have to manually trigger TTS, so you as a user choose what post that is run on. Is the thought here to provide a way for a user to limit the output of that generate button?

…e want. This value is passed through a filter so we can't blindly trust it but PHPStan was complaining about the old approach
@drzraf

drzraf commented Aug 27, 2026

Copy link
Copy Markdown

I guess I see no reason to complicate this further and introduce a setting to set the language when we can just rely on the LLM to match the language of the content given.

Assume the LLM fails the language detection for my post. What would be the steps to follow to overcome this (or would I be stuck, unable to associate a spoken version of the post?)

@dkotter

dkotter commented Aug 27, 2026

Copy link
Copy Markdown
Contributor Author

Assume the LLM fails the language detection for my post. What would be the steps to follow to overcome this (or would I be stuck, unable to associate a spoken version of the post?)

If this happens then yes, the audio generated wouldn't be what you want (or if the LLM errors instead, you wouldn't have any audio). This feels fairly unlikely to me that an LLM supports your language but can't generate speech without you telling it the language but I guess that may happen. Curious if this is a scenario you've run into before?

I personally still lean towards not complicating the UI by adding an additional setting a user has to consider and just assuming/hoping each AI service can detect the language properly. If we get actual user reports that run into problems with that, we can then look to add this in, though will need to figure out how to allow someone to set a language but not have that break integrations (like OpenAI) that don't support passing in a language.

@drzraf

drzraf commented Aug 27, 2026

Copy link
Copy Markdown

I'm completely with you about not complicating the UI.

What I'm a bit worried about is the limited internal API/hooks. Because if UI/network/LLM/service/quota/payment fails, that's what one would use (wp eval or whatever or read-out myself the article with the microphone) to resolve/workaround the problem. Simple UI but flexible API that also brings foundations for other plugins enhancements to build upon.

There are many ways language could be badly interpreted. For example, in STT (whisper), the understanding is language-neutral (because of the way training is done for natively mulilingual dataset/STT). Language acts as a "hint" but also condition the text output language. English with language=french would output the English transcribed text translated in French.

Non-English users may also have posts containing English expressions mixed up with their native language and I'm not sure these are situation where auto-detection is 100% reliable.

Basically, if install my local WordPress instance + AI endpoint + piper TTS + couple of custom voices, could I expect to manage language & LLM failure. Maybe that day I've to read-out myself the article using my microphone :)

Making language selection at the hook/filter level + audio attachment decoupled from LLM generation itself may make the whole workflow more resilient for unexpected cases. For a component relying so much on network/3rd-party service/non-deterministic process, it may be desirable.

superdav42 pushed a commit to WordPress/ai-provider-for-google that referenced this pull request Sep 3, 2026
## What?

Adds support for text to speech conversion

## Why?

We currently don't support text to speech conversion in this provider,
even though that is supported upstream in the PHP AI Client. By adding
this support, it allows others to build out text to speech systems using
this provider plugin.

## How?

- Introduce a new `GoogleTextToSpeechConversionModel` that handles all
requests to convert text to speech
- Ensure this model is loaded when a TTS generation request is made
- Ensure we properly map model options and capabilities when determining
what models support TTS

## Use of AI Tools

AI assistance: Yes
Tool(s): Claude Code
Model(s): Opus 4.8
Used for: Putting together a plan and executing on that plan. Plan
reviewed and modified by me and all code was reviewed and tested by me

## Testing Instructions

Hard to test on it's own as this plugin provides functionality but
doesn't actually do anything with that. Easiest approach is the
following:

1. Checkout the WordPress AI plugin from this
[PR](WordPress/ai#888)
2. Download this PR, activate and configure the Google Provider
3. In the AI settings page, turn on Text to Speech
4. Go to a post and find the Text to Speech panel in the sidebar
5. Click on the Generate Audio button and ensure it works as expected

## Changelog Entry

> Added - Support for text to speech conversion
@jeffpaul jeffpaul removed the [Status] Blocked Used to indicate unable to move forward label Sep 29, 2026
@jeffpaul jeffpaul mentioned this pull request Sep 29, 2026
2 of 24 tasks
@jeffpaul

Copy link
Copy Markdown
Member
  1. When exposing additional settings in the AI settings page, the Voice options appear but is an empty text field whereas I expected a dropdown.
Screenshot 2026-09-29 at 10 29 57 AM
  1. Potential a limitation of Playground (as this worked when I tested locally) but when using a basic 5 paragraph post I'm stuck constantly on the audio getting generated:
Screenshot 2026-09-29 at 10 34 06 AM

@dkotter

dkotter commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

When exposing additional settings in the AI settings page, the Voice options appear but is an empty text field whereas I expected a dropdown

This is expected though I agree it's a little jarring.

The problem here is the valid voice options are directly tied to the model being used. And we don't know that model here unless someone has specifically chosen that. In the scenario where someone has chosen a specific model, we do see if the request for that model comes with the valid voice options and if so, we do output those in select dropdown.

So under the right scenario, this will be a dropdown instead of a text input. But that "right scenario" isn't common, as neither Google nor OpenAI return the valid voice options with the model request so even if you choose a model for either of those, you still get a text input.

Really our only option would be to hardcode a list of voice options based on the model but that then becomes something for us to maintain and we won't be able to support every single model out there.

For now, if we want to expose this as a setting, I think an empty text input is the right choice. I've added some helper text below the input to make it more clear on what you need to do:

Screenshot 2026-09-29 at 4 13 39 PM

Potential a limitation of Playground (as this worked when I tested locally) but when using a basic 5 paragraph post I'm stuck constantly on the audio getting generated:

Yeah, in testing there it appears Playground disables WP-Cron, which we use to trigger these background processes. I imagine that's not a common scenario (cron completely disabled) but I guess worth a discussion on if we use something besides WP-Cron (Action Scheduler perhaps?).

I have adjusted things a bit though so if a job is never picked up, it now fails quicker (90 seconds vs 600 seconds). This gives enough time for the cron to actually fire. If after 90 seconds it still hasn't picked it up, we consider it a failure, we cancel the job and show an error message.

So if you test again in Playground and wait 90 seconds, you'll see processing stop and an error message show. We can lower that more if we want, though WordPress has a default 60 second lock on cron spawning so some times it may take around 60 seconds for the job to be picked up, so padding that a bit is where I got to 90 seconds.

@jeffpaul

jeffpaul commented Oct 2, 2026

Copy link
Copy Markdown
Member

The problem here is the valid voice options are directly tied to the model being used. And we don't know that model here unless someone has specifically chosen that. In the scenario where someone has chosen a specific model, we do see if the request for that model comes with the valid voice options and if so, we do output those in select dropdown.

Ok, then in this case if someone toggles on the developer tool for advanced options, let's also show the model dropdown selection before the voice dropdown and disable the voice dropdown until a model is selected (alternatively pre-select a model so that we can load in the voices in the dropdown and have that field enabled to start).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Needs review

Development

Successfully merging this pull request may close these issues.

4 participants