Expand backend lyrics support with richer sidecar formats and upgrade the
OpenSubsonic songLyrics implementation to the version 2 structured karaoke
contract, while preserving version 1 behavior by default.
Sidecar formats and parsing:
- Add a TTML parser (core/lyrics/ttml.go): clock time, offset time, bare
decimal seconds, nested timing contexts, and token-level <span> timing for
word/syllable karaoke. Parses Apple Music-style metadata tracks (translation
and pronunciation/transliteration) and agent metadata into per-track agents[]
plus per-cue-line agentId. Hydrates missing line timing from cue timing.
- Add an SRT parser (core/lyrics/srt.go).
- Add a LRCLIB Lyricsfile (.yaml/.yml) parser (model/lyricsfile.go): maps
per-word lines[].words[] to cues with inclusive UTF-8 byte offsets and
attributes overlapping lines to synthetic voice agents so parallel vocals
split correctly in the enhanced response.
- Extend LRC parsing for Enhanced LRC inline <mm:ss.xx> word-timing markers.
- Add UTF-8 BOM and UTF-16 LE support for TTML/LRC sidecars.
- Parse the above formats from embedded tags as well as sidecar files.
Source resolution:
- Default lyricspriority is now
".ttml,.yaml,.yml,.elrc,.lrc,.srt,.txt,embedded" so the new formats are
discoverable without manual configuration.
- Preserve configured source priority across duplicate media-file candidates
instead of only checking the first DB match, so higher-priority sidecar
lyrics on older duplicates can still win.
- Raise the embedded-lyrics tag maxLength to 1 MB to fit word-timed
TTML/Enhanced-LRC karaoke for a full song.
OpenSubsonic songLyrics v2:
- Advertise songLyrics versions [1, 2].
- With enhanced=true, getLyricsBySongId may return structuredLyrics.kind
(main/translation/pronunciation), cueLine[] line-level karaoke groupings,
cueLine.cue[] timed words/syllables with required UTF-8 byteStart/byteEnd,
reusable structuredLyrics.agents[], and cueLine.agentId references.
- Without enhanced=true, the response stays v1-compatible: no kind, no cueLine,
no agents, no non-main tracks; the existing line[] payload is always
populated so legacy clients keep working.
Contract details:
- cueLine is emitted only for synced lyrics with cue data.
- Within a cueLine, cue.end is normalized all-or-none and overlaps are removed;
overlaps across separate cueLines remain valid for parallel vocal layers.
- Missing cue end-times are filled from the next cue or the parent line.
- When cueLines share an index, the one whose agent has role "main" is first.
- LyricCue.Value is serialized as XML chardata; cues with nil start are skipped
rather than serialized as 0.
Refactoring:
- Move pure format parsers into model/ (lyrics.go, lyrics_ttml.go,
lyrics_srt.go, lyrics_embedded.go, lyricsfile.go) and extract Subsonic
response building into server/subsonic/lyrics.go.
- Centralize lyric-kind constants and add Lyrics.EffectiveKind/IsMainKind.
- Add gg.Clone helper.
Spec references:
https://github.com/opensubsonic/open-subsonic-api/discussions/213https://github.com/opensubsonic/open-subsonic-api/pull/218 (songLyrics v2)
https://github.com/opensubsonic/open-subsonic-api/pull/228 (cue byte offsets)