Conversation
…udio processing so we stay under API limits
…mbines those into a single file
…t to the AI Client to generate speech. Set up to be used for both Ability calls and background jobs
… Fires individual jobs, via cron, that will chunk content down and turn those into audio files, combining all files at the end
… a string of text or post content from a specific post ID
…nd imports it into the media library as an MP3 file
…he admin to make things look better. Add a better loading state and an icon to the button
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## develop #888 +/- ##
=============================================
- Coverage 81.53% 81.24% -0.30%
- Complexity 3071 3306 +235
=============================================
Files 129 138 +9
Lines 12253 13300 +1047
=============================================
+ Hits 9991 10806 +815
- Misses 2262 2494 +232
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
I would also suggest one missing feature: the ability to easily filter by post-type (or just pass the post to allow for more granular control [sticky/date/...]). The vast majority of websites would likely want to selectively use it based on specific post type. |
…ice no longer exists
…o it's used for each chunk
Add voice resolution and defaulting to the Text to Speech experiment
Thanks @saarnilauri! Made a few changes but looked good so I've merged that in now |
@drzraf I don't believe OpenAI supports passing in a language (or even Google), they just infer the language from the text you send. ElevenLabs does allow you to pass the language code but it also will default to using the language of the text you send. I guess I see no reason to complicate this further and introduce a setting to set the language when we can just rely on the LLM to match the language of the content given.
I guess not sure what the request is here? Right now you have to manually trigger TTS, so you as a user choose what post that is run on. Is the thought here to provide a way for a user to limit the output of that generate button? |
…e want. This value is passed through a filter so we can't blindly trust it but PHPStan was complaining about the old approach
Assume the LLM fails the language detection for my post. What would be the steps to follow to overcome this (or would I be stuck, unable to associate a spoken version of the post?) |
If this happens then yes, the audio generated wouldn't be what you want (or if the LLM errors instead, you wouldn't have any audio). This feels fairly unlikely to me that an LLM supports your language but can't generate speech without you telling it the language but I guess that may happen. Curious if this is a scenario you've run into before? I personally still lean towards not complicating the UI by adding an additional setting a user has to consider and just assuming/hoping each AI service can detect the language properly. If we get actual user reports that run into problems with that, we can then look to add this in, though will need to figure out how to allow someone to set a language but not have that break integrations (like OpenAI) that don't support passing in a language. |
|
I'm completely with you about not complicating the UI. What I'm a bit worried about is the limited internal API/hooks. Because if UI/network/LLM/service/quota/payment fails, that's what one would use ( There are many ways language could be badly interpreted. For example, in STT (whisper), the understanding is language-neutral (because of the way training is done for natively mulilingual dataset/STT). Language acts as a "hint" but also condition the text output language. English with language=french would output the English transcribed text translated in French. Non-English users may also have posts containing English expressions mixed up with their native language and I'm not sure these are situation where auto-detection is 100% reliable. Basically, if install my local WordPress instance + AI endpoint + Making language selection at the hook/filter level + audio attachment decoupled from LLM generation itself may make the whole workflow more resilient for unexpected cases. For a component relying so much on network/3rd-party service/non-deterministic process, it may be desirable. |
## What? Adds support for text to speech conversion ## Why? We currently don't support text to speech conversion in this provider, even though that is supported upstream in the PHP AI Client. By adding this support, it allows others to build out text to speech systems using this provider plugin. ## How? - Introduce a new `GoogleTextToSpeechConversionModel` that handles all requests to convert text to speech - Ensure this model is loaded when a TTS generation request is made - Ensure we properly map model options and capabilities when determining what models support TTS ## Use of AI Tools AI assistance: Yes Tool(s): Claude Code Model(s): Opus 4.8 Used for: Putting together a plan and executing on that plan. Plan reviewed and modified by me and all code was reviewed and tested by me ## Testing Instructions Hard to test on it's own as this plugin provides functionality but doesn't actually do anything with that. Easiest approach is the following: 1. Checkout the WordPress AI plugin from this [PR](WordPress/ai#888) 2. Download this PR, activate and configure the Google Provider 3. In the AI settings page, turn on Text to Speech 4. Go to a post and find the Text to Speech panel in the sidebar 5. Click on the Generate Audio button and ensure it works as expected ## Changelog Entry > Added - Support for text to speech conversion
…tale to account for failed CRON processing
…0 minutes for a job that was never picked up to get cancelled
Ok, then in this case if someone toggles on the developer tool for advanced options, let's also show the model dropdown selection before the voice dropdown and disable the voice dropdown until a model is selected (alternatively pre-select a model so that we can load in the voices in the dropdown and have that field enabled to start). |



What?
Adds a new experiment, Text to Speech, that allows editors the ability to on-demand generate speech for a post, including the title and post content. This is then displayed as a player on the front-end or that can be turned off on a post by post basis.
Why?
Adding support for Text to Speech allows sites to provide their users with the ability to listen to content instead of having to read the content.
How?
ai/speech-generationandai/speech-import. These can be used to generate speech from text and import that speech as an audio file into the Media LibraryUse of AI Tools
AI assistance: Yes
Tool(s): Claude Code
Model(s): Opus 4.8
Used for: Help with initial planning and implementation. Refinement, review and testing done by me
Testing Instructions
Screenshots
Changelog Entry