Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,14 +51,15 @@ A full-featured, self-hosted notes application with a block editor, an agentic A
- **Transitions** between segments — a dip through black or white, or a blend (dissolve, wipe, slide, circle open). A dip is drawn inside each segment and costs nothing; a blend needs the finished video encoded a second time, and the dialog says so
- **Motion on stills (Ken Burns)** — a slow zoom or pan over each image, with an adjustable travel distance and an option to include the title and chapter screens. Video clips are left alone, since the footage already moves. A drifting shot is rendered well above the output frame and scaled back down, which is what keeps the movement smooth instead of stepping a pixel at a time — it costs render time, so a shot that drifts is slower than one that doesn't. On a segment long enough that one sweep across it would be too slow to see, the motion cycles instead — drifting A to B, then B to A, in legs short enough to stay visible, rather than crawling once across the whole segment or holding still; a higher travel distance lengthens each leg. A per-segment encode timeout scales with the segment's own length rather than a flat cap, so a long section isn't cut off before it can finish
- **Background music** — an uploaded track mixed under the narration, ducking beneath speech and coming back up in the gaps, with its own level and fade in/out. A short track loops and a long one is cut to the video; the picture is never re-encoded to add it
- **Intro and outro clips** — an uploaded video played before the title screen and another after the last section, each whole and with its own sound, fitted to the frame. Nothing of the render is drawn over them (no watermark, overlay or waveform), the background music plays only between them, and each gets its own chapter marker
- **Quotes on screen** — a blockquote gets its own segment, with the words shown over the same picture while they are read, and a trailing "— name" line picked up as the attribution
- **Title screens**, optional chapter screens, chapter markers embedded in the MP4, an automatic thumbnail, and **subtitles** as an `.srt` sidecar, a track inside the MP4, or burned into the picture
- A chapter screen reads its own heading while the words are on screen; with chapter screens off, the heading is read inside the section it introduces
- An adjustable **pause at each heading**, held going in and coming out, so a section doesn't run straight into the next one — a full stop is all a voice has to separate them otherwise. Set it to zero to read headings on as ordinary prose
- An adjustable **pause at the end of every segment** — a paragraph, a section, a title or chapter screen — held after the last word before cutting to what's next, so a segment finishes rather than getting clipped by the cut
- **Every text size is adjustable** — title screen, chapter screen, watermark icon and caption, and the fixed overlay — set as a percentage of the frame height, so one choice holds at every resolution and aspect ratio
- Renders in the background with progress in the header and the browser tab, and can be cancelled mid-render — the finished video is attached to the note by the server, so it arrives even if you close the tab
- Choice of TTS voice and speaking rate per video, and options are remembered between renders — grouped into Format, Narration, Motion & audio, Branding and Structure tabs
- Choice of TTS voice and speaking rate per video, and options are remembered between renders — saved to your account, so they follow you to any browser, and grouped into Format, Narration, Motion & audio, Branding, Intro & outro and Structure tabs. Files the options point at (intro, outro, watermark icon, music) are kept in your media, and the unlinked-file sweep treats them as in use

### Sharing & export
- Export to PDF, Word (.docx), Markdown, HTML, MP3, MP4 video, or clipboard
Expand Down
47 changes: 45 additions & 2 deletions backend/app/asset_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,12 +22,12 @@
import os
import uuid
from dataclasses import dataclass
from typing import List, Optional, Tuple
from typing import Any, Iterator, List, Optional, Set, Tuple

import json
from sqlmodel import Session, col, or_, select

from app.models import Note, NoteAsset, Theme, TranscriptionJob, User, VideoRenderJob
from app.models import Note, NoteAsset, Theme, TranscriptionJob, User, UserSetting, VideoRenderJob
from app.routers.media import MEDIA_DIR, categorize_extension

logger = logging.getLogger(__name__)
Expand Down Expand Up @@ -269,6 +269,46 @@ def sync_note_assets(session: Session, note) -> int:
return 0


def _strings(value: Any) -> Iterator[str]:
"""Every string anywhere inside a decoded JSON value."""
if isinstance(value, str):
yield value
elif isinstance(value, dict):
for item in value.values():
yield from _strings(item)
elif isinstance(value, list):
for item in value:
yield from _strings(item)


def settings_media_filenames(session: Session, user_id: str) -> Set[str]:
"""Filenames in this user's media dir that one of their settings points at.

The video dialog's saved options hold the intro and outro clips, watermark,
music and background the user picked once and reuses on every render. Those
files belong to no note, so without this they would read as leaked and be
offered up for deletion by the unlinked-file sweep.
"""
names: Set[str] = set()
try:
values = session.exec(select(UserSetting.value).where(UserSetting.user_id == user_id)).all()
except Exception:
logger.exception("Could not read settings for user %s", user_id)
return names
for raw in values:
if not raw or MEDIA_URL_PREFIX not in raw:
continue
try:
decoded = json.loads(raw)
except (ValueError, TypeError):
continue
for text in _strings(decoded):
parsed = parse_media_url(text)
if parsed and parsed[0] == user_id:
names.add(parsed[1])
return names


def file_is_referenced(
session: Session,
user_id: str,
Expand Down Expand Up @@ -338,6 +378,9 @@ def file_is_referenced(
).first():
return True

if filename in settings_media_filenames(session, user_id):
return True

return False


Expand Down
5 changes: 5 additions & 0 deletions backend/app/routers/assets.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@
register_asset,
release_media_file,
remove_media_file,
settings_media_filenames,
sync_note_assets,
)
from app.database import get_session
Expand Down Expand Up @@ -327,6 +328,10 @@ def add_url(url: Optional[str]):
).all():
names.update(n for n in (result, subtitle, thumb) if n)

# Media a setting holds on to — the video dialog's intro and outro clips,
# watermark and music — is the user's own content, not a leak.
names.update(settings_media_filenames(session, user_id))

return names


Expand Down
17 changes: 16 additions & 1 deletion backend/app/video/ffmpeg.py
Original file line number Diff line number Diff line change
Expand Up @@ -659,6 +659,7 @@ def crossfade_total(durations: Sequence[float], overlap: float) -> float:
def build_music_command(
source: str, music: str, output: str,
*, duration: float, spec: MusicSpec, duck: bool,
start: float = 0.0, end: Optional[float] = None,
) -> List[str]:
"""Mix a background bed under the finished video, without touching the picture.

Expand All @@ -670,14 +671,28 @@ def build_music_command(
`-stream_loop -1` covers a bed shorter than the video and the output `-t`
truncates one that is longer. `normalize=0` stops `amix` halving the
narration to make room, which is the default and never what anyone wants.

`start`/`end` narrow the bed to a window of the video — the article between
an intro and an outro, which bring their own sound. The fades land at the
window's edges, and the bed is padded with silence past its end so the mix
and the ducking compressor run on to the end of the narration as before.
"""
fade_out_at = max(0.0, duration - max(0.0, spec.fade_out))
end = duration if end is None else max(0.0, min(end, duration))
start = max(0.0, min(start, end))
window = max(0.1, end - start)
fade_out_at = max(0.0, window - max(0.0, spec.fade_out))
bed = (f"[1:a]volume={max(0.0, min(1.0, spec.volume)):.3f},"
f"aresample=48000,aformat=channel_layouts=stereo")
if spec.fade_in > 0:
bed += f",afade=t=in:st=0:d={spec.fade_in:.3f}"
if spec.fade_out > 0:
bed += f",afade=t=out:st={fade_out_at:.3f}:d={spec.fade_out:.3f}"
if start > 0 or end < duration:
bed += f",atrim=end={window:.3f}"
if start > 0:
delay = int(round(start * 1000))
bed += f",adelay={delay}|{delay}"
bed += ",apad"
chains = [f"{bed}[m]"]

if duck:
Expand Down
22 changes: 22 additions & 0 deletions backend/app/video/options.py
Original file line number Diff line number Diff line change
Expand Up @@ -365,6 +365,23 @@ def _sane_scrim(cls, v: float) -> float:
return max(0.0, min(1.0, float(v)))


class BumperSpec(BaseModel):
"""A pre-made clip played before (intro) or after (outro) the article.

It plays whole, with its own sound, fitted to the frame like any other clip
in a note — but it is the user's own branding, so nothing of the render's is
drawn over it: no watermark, no text overlay, no waveform, and the music bed
stops short of it. Transitions still apply, so it joins the video the same
way every other segment does.
"""

enabled: bool = False
url: Optional[str] = None # /media/... video
# The uploaded file's own name. The file is stored under a UUID, so this is
# the only readable label the dialog has for it.
name: str = ""


class RenderOptions(BaseModel):
"""The complete render configuration. Persisted as JSON on the job row."""

Expand All @@ -391,6 +408,11 @@ class RenderOptions(BaseModel):
quotes: QuoteSpec = Field(default_factory=QuoteSpec)
code: CodeSpec = Field(default_factory=CodeSpec)

# Clips bracketing the whole video: the intro plays before the title
# screen, the outro after the last section. See BumperSpec.
intro: BumperSpec = Field(default_factory=BumperSpec)
outro: BumperSpec = Field(default_factory=BumperSpec)

# Append the finished video to the note as a playable block. Done by the
# worker rather than the browser so a render survives the tab being closed.
insert_into_note: bool = True
Expand Down
34 changes: 28 additions & 6 deletions backend/app/video/renderer.py
Original file line number Diff line number Diff line change
Expand Up @@ -287,12 +287,23 @@ def render(
if write_srt(os.path.join(work_dir, name), narration.cues):
shot_srt = name

if shot.chapter:
# An intro or outro's "chapter" names the clip, not a section of the
# article, so it must not become the line the overlay follows.
if shot.chapter and not shot.bumper:
current_chapter = shot.chapter

# An intro or outro is the user's own branded clip: nothing of the
# render's — watermark, overlay text, waveform — is drawn over it.
shot_options = options
if shot.bumper and options.waveform.enabled:
shot_options = options.model_copy(deep=True)
shot_options.waveform.enabled = False

shot_layer = layer
shot_overlay = overlay_name
if dynamic_overlay_text:
if shot.bumper:
shot_layer = shot_overlay = None
elif dynamic_overlay_text:
# The intro stretch before any heading is marked with the note's
# own title as its "chapter" (see segment()'s title card) — showing
# it again underneath the title would just repeat it, so that
Expand Down Expand Up @@ -349,7 +360,7 @@ def _encode(with_options: RenderOptions) -> None:
# A segment encoded by an earlier attempt that died later on (usually
# at the stitch) is reused rather than encoded again.
key = shot_cache.shot_key(
_argv(options),
_argv(shot_options),
[background, narration.path, shot_overlay, shot_srt],
work_dir,
)
Expand All @@ -358,14 +369,14 @@ def _encode(with_options: RenderOptions) -> None:
reused += 1
else:
try:
_encode(options)
_encode(shot_options)
except F.FFmpegError as exc:
# Losing a forty-minute render to one expensive segment is a bad
# trade when the two most expensive things in it are also the two
# least important. Retry once without them; the narration is
# already synthesised and cached, so this costs no speech.
logger.warning("Segment %d failed (%s) — retrying it plainer", index + 1, exc)
plain = options.model_copy(deep=True)
plain = shot_options.model_copy(deep=True)
plain.ken_burns.effect = "none"
plain.waveform.enabled = False
_encode(plain)
Expand Down Expand Up @@ -436,6 +447,14 @@ def _encode(with_options: RenderOptions) -> None:

final = "stitched.mp4"

# Where the article itself sits on the finished timeline. An intro or
# outro brings its own picture and sound, so the music bed and the poster
# frame are both taken from between them. The intro's last `overlap`
# seconds are already blending into the article; its end is the moment
# the blend has finished.
intro_end = durations[0] if shots[0].bumper == "intro" else 0.0
outro_length = durations[-1] if shots[-1].bumper == "outro" else 0.0

# ── background music ──────────────────────────────────────────────────
# A bed has to run continuously across shot boundaries, so it can only go
# on once the shots are joined. Mixing here re-encodes the audio alone —
Expand All @@ -452,6 +471,8 @@ def _encode(with_options: RenderOptions) -> None:
F.build_music_command(
final, music_path, "scored.mp4",
duration=scored_length, spec=options.music, duck=duck,
start=max(0.0, intro_end - overlap),
end=scored_length - outro_length,
),
cwd=work_dir, timeout=1800,
)
Expand Down Expand Up @@ -498,7 +519,8 @@ def _encode(with_options: RenderOptions) -> None:
F.build_poster_command(
os.path.join(user_dir, video_filename),
os.path.join(user_dir, thumbnail_filename),
at_seconds=min(1.0, max(0.0, timeline / 2)),
at_seconds=intro_end + min(
1.0, max(0.0, (timeline - intro_end - outro_length) / 2)),
),
timeout=120,
)
Expand Down
35 changes: 35 additions & 0 deletions backend/app/video/segmenter.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,9 @@ class Shot:
# `narration` above only carries its text when narrate_code is on — see
# segment()'s codeBlock branch.
code_text: Optional[str] = None
# "intro" or "outro" for a clip bracketing the article (see _bumper_shot).
# The renderer draws nothing of its own over one of these.
bumper: Optional[str] = None
# Set on the second half of a sounded-clip pair, purely for readable logs.
label: str = ""

Expand Down Expand Up @@ -561,4 +564,36 @@ def flush(next_shot: Optional[Shot]) -> None:
continue
trimmed.append(shot)
result.shots = trimmed

# The intro and outro bracket the article rather than stand in for it, so a
# note with nothing of its own to show still renders nothing.
if result.shots:
intro = _bumper_shot("intro", options, user_id, media_dir, result.warnings)
outro = _bumper_shot("outro", options, user_id, media_dir, result.warnings)
if intro is not None:
result.shots.insert(0, intro)
if outro is not None:
result.shots.append(outro)
return result


def _bumper_shot(
which: str, options: RenderOptions, user_id: str, media_dir: str, warnings: List[str],
) -> Optional[Shot]:
"""The intro or outro clip as a shot, or None when it is off or unusable.

It is a sounded clip like any other in a note: played whole, with its own
audio when it has some and silence when it doesn't, which the renderer
probes for itself.
"""
spec = options.intro if which == "intro" else options.outro
if not spec.enabled or not spec.url:
return None
path = resolve_media_path(spec.url, user_id, media_dir)
if path is None or _media_kind(spec.url) != "video":
warnings.append(f"The {which} was skipped: that clip could not be read.")
return None
return Shot(
kind="video_sound", background=path,
chapter=which.capitalize(), bumper=which, label=which,
)
45 changes: 44 additions & 1 deletion backend/tests/test_note_assets.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@
release_media_file,
sync_note_assets,
)
from app.models import Note, NoteAsset, Theme, User
from app.models import Note, NoteAsset, Theme, User, UserSetting
from app.routers.assets import _ai_eligible, _role_for

USER = "user-1"
Expand Down Expand Up @@ -351,6 +351,49 @@ def test_release_keeps_a_file_used_as_a_theme_background(session, media_dir):
assert release_media_file(session, USER, url) is False


def _save_setting(session, key, value, user_id=USER):
session.add(UserSetting(user_id=user_id, key=key, value=json.dumps(value)))
session.commit()


def test_release_keeps_a_file_a_saved_setting_points_at(session, media_dir):
"""The video dialog's intro clip belongs to no note, but it is still in use."""
url = media_file(media_dir, USER, "intro.mp4")
_save_setting(session, "video_render_options",
{"intro": {"enabled": True, "url": url}, "music": {"url": None}})

assert release_media_file(session, USER, url) is False
assert (media_dir / USER / "intro.mp4").exists()


def test_settings_media_filenames_reads_nested_values_and_only_the_users_own(session):
_save_setting(session, "video_render_options", {
"intro": {"url": f"/media/{USER}/intro.mp4"},
"outro": {"url": f"/media/{USER}/outro.mp4"},
"watermark": {"url": f"/media/{OTHER_USER}/theirs.png"},
"diagram_images": {"b1": f"/media/{USER}/diagram.png"},
"list": [f"/media/{USER}/in-a-list.mp3", "not a url", 3, None],
})
_save_setting(session, "video_render_options", {"intro": {"url": f"/media/{OTHER_USER}/x.mp4"}},
user_id=OTHER_USER)
session.add(UserSetting(user_id=USER, key="broken", value="/media/{not json"))
session.commit()

assert asset_utils.settings_media_filenames(session, USER) == {
"intro.mp4", "outro.mp4", "diagram.png", "in-a-list.mp3",
}


def test_the_unlinked_sweep_does_not_offer_up_a_settings_file(session, media_dir):
from app.routers.assets import _referenced_filenames

url = media_file(media_dir, USER, "outro.mp4")
_save_setting(session, "video_render_options", {"outro": {"enabled": False, "url": url}})

# Switched off but still chosen: the clip is kept for when it's switched back on.
assert "outro.mp4" in _referenced_filenames(session, USER)


def test_release_never_touches_another_users_file(session, media_dir):
"""Content pasted from a shared note points into the author's media dir, not ours."""
url = media_file(media_dir, OTHER_USER, "theirs.png")
Expand Down
Loading
Loading