# Stage 9 — AI Narration with ElevenLabs

## Status

Implemented locally as of 2026-10-04. Rollout remains protected by `ELEVENLABS_TTS_ENABLED=false` until the additive migration, production credentials, queue-worker restart, and one-slide Arabic smoke test are completed.

Stage 9 belongs to the `main` product line. It completes the text-to-speech work that Stage 3 deliberately deferred, while preserving the existing PowerPoint-to-SCORM audio, media, revision, and package contracts.

The implementation must use additive migrations. Existing projects, slides, uploaded narration, other audio, media settings, generated packages, and private files must remain valid and unchanged.

## Summary

Store one editable narration text for every course slide, show that same text in both **Slides Structure** and **Files & Media**, and let an authorized author generate Arabic narration through either an activated personal ElevenLabs account or the application fallback account.

ElevenLabs returns MP3 audio. The generated file then becomes a normal application audio media item and uses the existing audio processing, settings, preview, and SCORM packaging workflow. The only behavioral difference is that authors can see that its origin is **AI generated · ElevenLabs**.

## Agreed Stage 9 Decisions

- Every user can securely store a personal ElevenLabs API key and voice ID and explicitly activate or deactivate them.
- An activated and complete personal account is the primary source for narration requested by that user. If it is inactive or either its API key or voice ID is missing, the application account is used automatically.
- Personal API keys are encrypted with Laravel's application key. No API key is ever returned to the browser, placed in a job payload or provenance, logged, or included in a SCORM package.
- The production model is the quality-focused `eleven_v4`; `eleven_flash_v2_5` remains the lower-cost fallback when budget matters more than maximum expressiveness.
- The initial output format is `mp3_44100_128`.
- The model and output format remain application-managed. Each user may configure only their API key, voice ID, and activation state.
- Every course slide owns one narration-text record, including slides whose text is initially empty.
- Speaker notes are the preferred initial narration source.
- Visible PowerPoint slide text is used as a draft when speaker notes are absent.
- Authors can edit the narration before spending ElevenLabs credits.
- **Slides Structure** and **Files & Media** edit the same narration record, not separate copies.
- Audio generation is asynchronous, with one queued generation job per slide.
- Generated MP3 files become normal `audio` media items and support every existing audio setting.
- AI provenance is authoring metadata. It is not shown to learners in the exported SCORM course.
- Changing narration text never regenerates automatically. It marks existing generated audio as outdated.
- Existing manually uploaded audio is never overwritten by narration generation.

## Goals

- Give every slide a persistent, editable narration script.
- Import a useful narration draft from the source PowerPoint when possible.
- Make narration easy to review in both slide organization and media-management workflows.
- Generate one slide, selected slides, all missing slides, or all outdated slides.
- Preserve the last usable audio until replacement generation succeeds.
- Reuse the existing audio player, overall audio settings, media worker, preview, and SCORM builder.
- Track enough AI provenance to identify, audit, and safely regenerate generated narration.
- Keep the workflow bilingual, responsive, keyboard accessible, and correct in Arabic RTL.

## Non-goals

- User-selectable TTS models in the initial release.
- Voice cloning or voice training.
- Multiple narration languages for one slide.
- Automatic translation, dubbing, or transcript generation from existing audio.
- Automatic generation immediately after PowerPoint upload or text editing.
- OCR of rendered slide images.
- Reading quiz questions, answers, media transcripts, or downloadable resources as narration.
- A second audio player or a separate SCORM narration format.
- Learner-time calls to ElevenLabs. Exported courses must remain self-contained.
- Self-hosting ElevenLabs models.

## 1. Narration Text Source

### Source priority

For each source PowerPoint slide, create the initial narration draft using this order:

1. Speaker notes, when meaningful notes exist.
2. Visible slide text, when speaker notes are empty.
3. Empty narration text when neither source is available.

Speaker notes take priority because they are normally written to be spoken. Visible slide text is only a starting draft; bullet points and headings may not form natural narration and must remain editable before generation.

### PowerPoint extraction

Add a dedicated narration extractor for `.pptx` files. It must:

- Read slide order through `ppt/presentation.xml` and its relationships rather than assuming ZIP filename order.
- Resolve each slide's notes relationship safely.
- Extract meaningful text from notes while excluding slide-number, date, footer, and other presentation placeholders.
- Extract ordinary visible DrawingML text when no meaningful notes exist.
- Preserve paragraph boundaries and normalize repeated whitespace without joining unrelated words.
- Return text keyed by the original source slide number.
- Use the same ZIP entry-count, expanded-size, unsafe-path, and XML protections as the existing PowerPoint inspection pipeline.
- Never execute macros, relationships, embedded objects, or external links.
- Never use OCR in Stage 9.

Text extraction is best effort. Unsupported diagrams, equations, images of text, and unusual third-party PowerPoint structures may produce an empty or incomplete draft; the user can correct it manually.

### Existing and manually created slides

- New PowerPoint conversions receive extracted drafts when their course-slide records are created.
- Existing projects receive a one-time, explicit or lazy import from their retained source PowerPoint when Stage 9 narration is first opened.
- Import fills only empty narration records. It must never replace text already edited by a user.
- Blank and uploaded-image slides begin with empty narration text.
- Duplicating a slide copies its narration text as a new editable draft but does not copy or share its generated audio.
- Reordering a slide keeps narration connected through the stable `course_slide_id`.
- Removing a slide retains its narration and generated media so restoration is lossless.

## 2. Narration Data Model

Add a `conversion_slide_narrations` table with one row per course slide:

- `course_slide_id`, UUID primary key and foreign key to `conversion_course_slides`, cascade on permanent deletion.
- `conversion_id`, UUID foreign key to `conversions`, cascade on project deletion.
- `text`, nullable long text.
- `source`, short string: `speaker_notes`, `slide_text`, `manual`, `copied`, or `empty`.
- `revision`, unsigned integer used for narration-specific optimistic concurrency.
- `text_hash`, nullable SHA-256 of the normalized current text.
- `generated_text_hash`, nullable SHA-256 of the text used by the last successful generation.
- `status`: `empty`, `draft`, `queued`, `generating`, `ready`, `outdated`, or `failed`.
- `media_item_id`, nullable UUID reference to the active generated audio item.
- `last_error_code` and `last_error_message`, nullable and safe for author display.
- `generated_at`, nullable timestamp.
- Created and updated timestamps.

Required constraints and indexes:

- Unique ownership through the primary `course_slide_id`.
- Index on `(conversion_id, status)` for the narration workspace and batch actions.
- The referenced media item must belong to the same conversion and slide; this ownership check must also be enforced in application code.

Do not store the ElevenLabs API key, raw provider credentials, or sensitive HTTP responses in this table.

Extend `users` additively with `elevenlabs_enabled`, encrypted `elevenlabs_api_key`, and nullable `elevenlabs_voice_id`. The API returns only activation/configuration booleans and the non-secret voice ID. A blank key update preserves the stored key; an explicit removal action clears it and deactivates the personal integration. These encrypted values depend on the deployment `APP_KEY`, which must remain stable and securely backed up.

### AI media provenance

Extend `conversion_media_assets` additively with nullable provenance fields:

- `origin`, where Stage 9 writes `ai_tts`; existing `NULL` rows continue to mean ordinary uploaded media.
- `provenance`, nullable JSON containing the provider, model, voice ID, output format, credential source (`user` or `application`), narration text hash, and provider request identifier when available.

The authoring UI displays **AI generated · ElevenLabs** when `origin=ai_tts` and `provenance.provider=elevenlabs`.

Do not place the API key, full request payload, or narration text inside provenance JSON. The canonical text remains in `conversion_slide_narrations`; the normal media `transcript` stores the successful text snapshot for authoring and regeneration, but AI narration transcripts are not displayed in the learner preview or packaged SCORM.

## 3. Narration States

| State | Meaning |
| --- | --- |
| Empty | No usable narration text exists; Generate is disabled. |
| Draft | Text exists but no successful AI narration exists. |
| Queued | A generation request is waiting for the queue worker. |
| Generating | The worker is currently calling ElevenLabs or storing the result. |
| Ready | The active generated audio matches the current text hash. |
| Outdated | Text changed after the last successful generation. The old audio remains usable. |
| Failed | The latest generation failed. The previous ready audio remains usable when one exists. |

Status changes must be transactional and must not leave a slide pointing to a partially written asset.

## 4. Slides Structure Experience

Show a narration section beneath every active slide:

- Editable multiline narration text.
- Source label such as **Imported from speaker notes**, **Imported from slide text**, or **Edited manually**.
- Save state: saving, saved, failed, or changed in another session.
- Audio state: not generated, generating, ready, outdated, or failed.
- A link or action to open the same slide in Files & Media for generation and audio configuration.

Narration editing uses a short debounce or save-on-blur and the narration revision. A `409` conflict must preserve the user's local text and offer reload/retry; it must never silently replace another editor's saved narration.

The structure workspace does not need to duplicate all audio-player settings. Its purpose is rapid slide-by-slide script review and editing.

## 5. Files & Media Experience

Add an **AI Narration** view or section inside the established Files & Media workspace. Each active slide card shows:

- Slide thumbnail, number, and title.
- The same editable narration text used by Slides Structure.
- Character count.
- Narration and audio status.
- **Generate audio** when text exists and no current generation is running.
- **Preview audio** after a successful generation.
- **Regenerate** when audio is ready or outdated.
- **Retry** after a safe failure.
- **Delete generated audio**, which preserves the narration text.
- The **AI generated · ElevenLabs** badge on generated audio.

Workspace-level actions:

- Generate selected slides.
- Generate all missing narration.
- Regenerate all outdated narration.
- Show progress such as `12 of 30 slides ready`.
- Filter by empty, draft, generating, ready, outdated, or failed.

The narration text must always be visible and editable before generation. Text editing in either workspace updates the same stored record and becomes visible in the other workspace after save/refetch.

## 6. ElevenLabs Application Configuration

Use backend environment configuration for Stage 9:

```dotenv
ELEVENLABS_TTS_ENABLED=false
ELEVENLABS_API_KEY=
ELEVENLABS_MODEL_ID=eleven_v4
ELEVENLABS_VOICE_ID=
ELEVENLABS_OUTPUT_FORMAT=mp3_44100_128
ELEVENLABS_CONNECT_TIMEOUT_SECONDS=10
ELEVENLABS_REQUEST_TIMEOUT_SECONDS=120
ELEVENLABS_MAX_CONCURRENT_REQUESTS=2
```

Requirements:

- Configuration must be read through Laravel config, never by calling `env()` from application services.
- The key must be a restricted server-side ElevenLabs key with only required permissions and an appropriate credit limit.
- The key must never be logged, serialized into a job payload, returned by an API, or included in frontend configuration.
- The configured model and voice are application-wide defaults.
- The authoring UI may display the configured voice name and service availability but not the secret.
- When the feature is disabled or incomplete, generation controls show a clear unavailable state while narration editing remains usable.

An administrator-only connection check may validate the configured credentials and voice without exposing them. It must use a very short test phrase and clearly warn that the check consumes credits.

## 7. Generation Workflow

1. The author edits and saves narration text.
2. The author requests generation for one or more slides.
3. Laravel validates edit permission, active slide ownership, non-empty text, revision, length, and feature configuration.
4. The application assigns a generation token to the saved narration revision.
5. One `GenerateSlideNarration` job is dispatched per slide to the existing `media` queue. Its payload contains only the narration ID and generation token.
6. The job reloads text and server credentials, then rechecks the token, text hash, slide state, and project ownership before spending credits.
7. The job calls the ElevenLabs Text-to-Speech API from Laravel.
8. The returned MP3 is written to a new private media-asset source path.
9. The existing `MediaProcessor` normalizes the pending source through FFmpeg before it can become active.
10. In one database transaction, the application creates or updates the standard audio media item, records AI provenance, sets the transcript, swaps the normalized asset, and updates narration status.
11. Only after processing and the guarded swap succeed does narration become ready; the course package is then marked outdated if a package already exists.
12. The next SCORM build packages the MP3 through the existing media path.

### Media-item behavior

- The first successful generation creates one ordinary `ConversionMediaItem` with `type=audio`, attached to the narration's `course_slide_id`.
- The new item initially adopts the course's overall audio settings.
- Authors may then apply every existing audio setting: layer, placement, player display, playback trigger, visibility, controls, completion, transcript, and overall/custom settings.
- Regeneration preserves the existing media item's ID, tracking key, position, and settings; it replaces only its asset after the new audio is safely available.
- Other uploaded or generated audio on the slide is not changed.
- The old generated asset remains active until the replacement has been stored and processed successfully.
- A manual file replacement through the normal media interface changes the asset origin to uploaded media and removes the AI badge. The narration text remains as a draft and is no longer considered matched to that manually uploaded file.
- Moving the generated media item retains AI provenance but does not change the narration owner automatically. The implementation must either move the narration association in the same guarded transaction or block a move that would separate them; silently mismatching text and slide is prohibited.

### Text length

Stage 9 accepts a maximum of 8,000 Unicode characters per slide, below the selected model's current 10,000-character request limit. Empty or whitespace-only narration is rejected.

Stage 9 does not automatically split one slide into multiple ElevenLabs requests. An over-limit slide receives an actionable validation message so the author can shorten or divide the narration deliberately.

## 8. Queue, Retry, and Credit Safety

- Generation is always asynchronous; HTTP authoring requests never wait for audio generation.
- A course batch dispatches independent slide jobs so one failure does not cancel successful slides.
- Limit concurrent ElevenLabs calls to the configured application maximum.
- Retry explicit rate limits and temporary provider failures with bounded exponential backoff and jitter.
- Do not blindly retry an ambiguous timeout after the provider may have completed and charged the request. Mark it failed for author review unless the provider supplies a safe request identifier or idempotency mechanism.
- Before calling ElevenLabs, skip generation when ready audio already matches the same text hash, model, voice, and output format.
- A stale queued job must stop before spending credits when narration text or revision changed after dispatch.
- Store only safe provider error codes/messages. Never store or display response headers containing credentials or internal details.
- Cancelling or removing a course does not allow a late job to recreate deleted rows or files.

Stage 9 uses the existing `media` worker to avoid adding another Supervisor program. A future high-volume release may introduce a dedicated TTS queue after real throughput is measured.

## 9. API Contract

Add narration endpoints under the existing authenticated conversion routes:

- `GET /api/v1/powerpoint-to-scorm/conversions/{conversion}/narrations`
- `PATCH /api/v1/powerpoint-to-scorm/conversions/{conversion}/slides/{slide}/narration`
- `POST /api/v1/powerpoint-to-scorm/conversions/{conversion}/narrations/generate`
- `POST /api/v1/powerpoint-to-scorm/conversions/{conversion}/slides/{slide}/narration/retry`
- `DELETE /api/v1/powerpoint-to-scorm/conversions/{conversion}/slides/{slide}/narration/audio`

The collection response includes slide identity, text, source, revision, status, current/generated hashes, safe error information, and the ordinary media item when one is attached.

The batch-generation request accepts explicit `course_slide_ids` or a reviewed mode such as `missing` or `outdated`. It returns `202 Accepted` with batch counts; clients poll the normal narration document or use the application's existing progress mechanism.

Every endpoint must:

- Use the central conversion authorization rules.
- Verify that nested slide and media IDs belong to the route conversion.
- Reject removed or excluded slides for new generation while retaining their existing data.
- Use stable error codes in addition to localized messages.
- Preserve existing media and slide response fields.

Suggested stable errors:

- `tts_disabled`
- `tts_not_configured`
- `narration_empty`
- `narration_too_long`
- `narration_revision_conflict`
- `narration_already_generating`
- `tts_rate_limited`
- `tts_insufficient_credits`
- `tts_provider_unavailable`
- `tts_generation_failed`

## 10. Security, Privacy, Licensing, and Cost

- Use an ElevenLabs paid plan that permits commercial use before production generation.
- Confirm rights and consent for the configured voice. Stage 9 does not create or train cloned voices.
- Narration text is sent to ElevenLabs. This external processing must be disclosed in the application's privacy documentation.
- Default ElevenLabs retention and the account's data-use settings must be reviewed before sending confidential course material.
- If contractual zero retention is required, resolve the appropriate ElevenLabs Enterprise arrangement before production use.
- Enforce project edit authorization before generation because generation consumes shared application credits.
- Record internal usage by user, course, slide, character count, success/failure, model, and timestamp without storing the API key.
- Add configurable per-project or per-user rate/credit safeguards if abuse becomes possible; the initial release must at least rate-limit generation endpoints and prevent duplicate same-text jobs.

## 11. SCORM Behavior

- Generated narration is stored and packaged as ordinary MP3 audio.
- The exported course never calls ElevenLabs and requires no ElevenLabs key or internet connection for narration playback.
- All existing audio presentation and completion settings apply.
- The successful narration text remains available in authoring but is suppressed in the learner full preview and omitted from packaged SCORM configuration.
- The AI badge and provider metadata remain authoring-only and are not included in learner-visible UI.
- Generating, regenerating, deleting, moving, or manually replacing narration audio increments the relevant media revision and marks an existing package outdated.
- A failed or in-progress replacement never removes the last valid packaged audio source.

## 12. Migration and Existing-Project Safety

- Add only the narration table, nullable media-provenance columns, indexes, and foreign keys required by Stage 9.
- Do not rewrite existing course slides, media items, assets, tracking keys, paths, settings, revisions, or generated packages during migration.
- Do not call ElevenLabs, extract every stored PowerPoint, or generate audio inside a migration.
- Existing media with null provenance remains ordinary uploaded media.
- Narration rows for existing slides may be created lazily or by an idempotent application task after deployment.
- Text extraction fills only empty records and is safe to retry.
- Feature-disabled deployments must continue operating exactly as before, while the additive schema remains harmless.

## 13. Deployment

1. Confirm a paid ElevenLabs account, commercial-use terms, voice rights, credit limit, and restricted API key.
2. Back up the production database, `.env`, and `storage/app/private` and verify the backups are readable.
3. Deploy backend and frontend code with `ELEVENLABS_TTS_ENABLED=false`.
4. Run the additive Stage 9 migration.
5. Configure the API key, model, voice, output format, timeouts, and concurrency limit in the production environment.
6. Clear Laravel configuration caches and restart the existing `docudeck-media-worker` so it receives the new configuration and job class.
7. Verify narration text import/editing without enabling generation.
8. Enable ElevenLabs TTS and clear configuration cache again.
9. Generate one short Arabic test narration in a non-production course and verify MP3 processing, preview, media settings, package invalidation, SCORM regeneration, and Moodle playback.
10. Enable author access only after the smoke test passes.

Never during deployment:

- Put the production API key in source control, frontend environment variables, logs, screenshots, or documentation.
- Replace production `.env` with `.env.example`.
- Generate a new Laravel `APP_KEY`.
- Run `migrate:fresh`, `db:wipe`, truncation, or destructive backfills.
- Remove existing audio or regenerate every course automatically.

Rollback begins by setting `ELEVENLABS_TTS_ENABLED=false`, clearing configuration cache, and restarting workers. Leave additive narration data and provenance columns in place while corrected code is deployed; do not drop real narration text or generation history during an ordinary code rollback.

## 14. Test Plan

### PowerPoint extraction

- Speaker notes take priority over visible slide text.
- Slides without notes use visible text.
- Slides without either source remain empty.
- Notes placeholders, slide numbers, dates, and footers are excluded.
- Arabic, English, mixed Arabic/English, punctuation, and line breaks remain readable.
- Unsafe ZIP paths, oversized expansion, malformed XML, and external entities remain rejected.
- Existing projects import only into empty narration records.

### Narration editing

- Every active slide appears in both workspaces with the same text.
- Editing in either workspace updates the shared record.
- Revision conflicts preserve local text and return `409`.
- Reordering keeps narration attached to the stable slide.
- Duplicate, remove, restore, exclude, and permanently delete behavior matches the agreed rules.
- Empty and over-limit narration cannot be generated.

### Generation

- Tests fake ElevenLabs HTTP responses and never spend real credits.
- Validate request URL, authentication header handling, model, voice, output format, and exact saved text snapshot.
- One slide and batch generation dispatch the expected jobs.
- Same-text duplicate requests do not consume duplicate jobs or credits.
- Stale jobs stop before calling the provider.
- Rate limits, insufficient credits, provider errors, malformed audio, and ambiguous timeouts produce safe recoverable states.
- A failed regeneration retains the previous playable audio.
- Successful generation creates an ordinary audio item, transcript, provenance, and processing job.

### Media and SCORM regression

- Generated audio receives overall audio settings initially.
- Every existing audio setting works on generated narration.
- Regeneration preserves the media tracking key and settings.
- Existing manually uploaded audio is not replaced.
- Manual replacement removes the AI badge as agreed.
- The SCORM ZIP includes the normalized `audio.mp3` and no API credentials or provider calls.
- Learner preview and Moodle SCORM 1.2 playback work offline after package download.

### User experience

- Desktop and mobile layouts remain usable with long Arabic narration.
- English and Arabic localization and RTL direction are correct.
- Keyboard focus, labels, status announcements, errors, progress, and buttons are accessible.
- Viewer access remains read-only and sends no generation or narration-save requests.
- AI badges are visible to authors and absent from learner output.

## 15. Acceptance Criteria

Stage 9 is complete only when:

- Every course slide has one persistent narration record.
- Speaker notes or visible slide text create a safe editable initial draft where available.
- Slides Structure and Files & Media show and edit the same narration text.
- Authors always review text before generation and text edits never spend credits automatically.
- One-slide and batch ElevenLabs generation run asynchronously using the requesting user's activated credentials first and the application configuration as fallback.
- Successful MP3 output becomes a normal audio media item with all existing settings.
- Generated audio is clearly flagged as AI-generated in authoring interfaces.
- Existing uploaded audio and old projects remain unchanged.
- Regeneration preserves the last valid audio until the replacement succeeds.
- SCORM packages contain local MP3 files and no ElevenLabs secret or runtime dependency.
- Commercial licensing, voice rights, privacy, retention, and credit controls have been approved for production.
- Backend, frontend, queue, extraction, media, SCORM, Arabic/RTL, mobile, and Moodle acceptance tests pass.

## Official ElevenLabs References

- Text-to-Speech API: <https://elevenlabs.io/docs/api-reference/text-to-speech/convert>
- Models and limits: <https://elevenlabs.io/docs/overview/models>
- API-key security: <https://elevenlabs.io/docs/overview/administration/workspaces/api-keys>
- Pricing: <https://elevenlabs.io/pricing/api>
- Commercial-use guidance: <https://help.elevenlabs.io/hc/en-us/articles/13313564601361-Can-I-publish-the-content-I-generate-on-the-platform>
- Zero Retention Mode: <https://elevenlabs.io/docs/eleven-api/resources/zero-retention-mode>

## Assumptions and Defaults

- The application account remains an optional fallback; users use their own ElevenLabs account only when the personal integration is activated and both its API key and voice ID are present.
- `eleven_v4` is the production default because it is ElevenLabs' highest-quality, most expressive model and supports Arabic among 90+ languages. Before broad rollout, compare representative Arabic and mixed-English slides and confirm the higher generation cost is acceptable. Switch the environment value to `eleven_flash_v2_5` when lower API cost is preferred; no code or database change is required.
- Each user can activate one personal voice ID; otherwise the application-level voice is used.
- MP3 at 44.1 kHz and 128 kbps is the canonical provider response.
- Narration generation is optional; authors can keep narration text without generating audio.
- New generated audio inherits overall audio settings.
- Narration text is the canonical source for regeneration and the successful snapshot becomes the audio transcript.
- AI provenance is visible to authors only.
- Existing `media` workers process both TTS-generation and audio-normalization jobs in the initial release.
