mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 12:34:43 -04:00
* feat(schema): validate portable speaker profiles Add the versioned profile schema for explicit speaker enrollment. Validate compatibility against separately supplied loaded-encoder metadata. Reject unusable speakers, invalid vectors, and inconsistent clean spans. This slice does not change HTTP routes, backend integration, or the UI. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet): export profiles with transcripts Export opt-in speaker profiles and trusted encoder metadata. Replay registrations by ID so duplicate display names keep independent vectors. Use one profile-capable diarization for slots, names, and clean spans. Assign timestamped ASR words to those slots without a second diarization. Preserve legacy opt-out and no-ASR behavior, and propagate failures. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio): enroll portable speaker profiles Gate profile exports with voice-recognition permission and validate registration against metadata from the loaded encoder. Preserve audio enrollment and independent registrations with duplicate display names. Exclude diarization and registration exchanges before API trace capture so persisted traces cannot retain profile vectors or JSON audio. Defer candidate dimensions to trusted loaded metadata. Sort candidates by registration ID so incompatible profiles cannot suppress legacy voices through registry iteration order. Keep portable identity checks closed when trusted metadata is unavailable. Test persisted traces, explicit slot zero, and selection through offline and live transport. Document privacy and the ephemeral registry lifecycle. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(ui): remember speakers from diarization Add a Studio page for diarization and opt-in speaker profiles. Preview clean intervals from the original recording before explicit registration. Join profiles by raw speaker labels, preserve duplicate names, and relabel turns only after a successful save. Discard stale results when the model or recording changes. Share registration metadata with voice management without storing vectors or recordings from this flow. Document permissions and the global, ephemeral registry. Cover enrollment, permissions, previews, and asynchronous races with mocked Playwright tests. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: clarify HTTP speaker enrollment support Replace the stale enrollment limitation with the current HTTP workflow. Distinguish native transport from explicit registration and link its docs. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet): pin merged speaker profile support Use the merged commit from mudler/parakeet.cpp#80. Its tree matches the previously accepted native pin. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add diarization enrollment setup example Connect the existing gallery modes to the speaker enrollment workflow. Show installation, private profile export, explicit raw-slot registration, and later recognition without another export. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): explain diarization speaker profiles Put the diarization walkthrough on the LocalAI website in the feature PR. Cover the three gallery modes, explicit enrollment, and privacy limits. Link setup instructions and keep availability conditional on feature support. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): focus diarization on everyday use Explain what users can do with recordings before the setup steps. Replace the technical walkthrough with a short Studio guide and link readers to the existing reference for model names and developer use. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): lead with speaker capabilities Present speaker recognition through everyday uses and a short UI flow. Keep technical reference details in the existing documentation. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(diarization): satisfy Go lint checks Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
56 lines
2.5 KiB
Go
56 lines
2.5 KiB
Go
package schema
|
|
|
|
// DiarizationSegment is one continuous span of speech attributed to a
|
|
// single speaker. Times are in seconds. Speaker is the normalized label
|
|
// (SPEAKER_NN, zero-padded, stable across segments); Label preserves the
|
|
// raw backend-emitted identifier for clients that already track their
|
|
// own speaker dictionary.
|
|
type DiarizationSegment struct {
|
|
Id int `json:"id"`
|
|
Speaker string `json:"speaker"`
|
|
Label string `json:"label,omitempty"`
|
|
Start float64 `json:"start"`
|
|
End float64 `json:"end"`
|
|
Text string `json:"text,omitempty"`
|
|
// Name is the registered speaker this segment was matched to, and NameScore
|
|
// the cosine similarity of the match. Both are omitted when the backend did
|
|
// not identify the speaker. Speaker stays the normalized SPEAKER_NN label.
|
|
Name string `json:"name,omitempty"`
|
|
NameScore float32 `json:"name_score,omitempty"`
|
|
}
|
|
|
|
// DiarizationSpeaker summarizes one speaker across the whole audio so
|
|
// clients can build per-speaker UIs (timeline strips, talk-time charts)
|
|
// without re-aggregating the segment list.
|
|
type DiarizationSpeaker struct {
|
|
Id string `json:"id"`
|
|
Label string `json:"label,omitempty"`
|
|
Name string `json:"name,omitempty"`
|
|
TotalSpeechDuration float64 `json:"total_speech_duration"`
|
|
SegmentCount int `json:"segment_count"`
|
|
}
|
|
|
|
// DiarizationResult is the JSON payload returned by /v1/audio/diarization.
|
|
// Speakers and segment text are omitted when empty so the default `json`
|
|
// response stays minimal; verbose_json keeps both populated.
|
|
type DiarizationResult struct {
|
|
SpeakerProfiles *SpeakerProfiles `json:"speaker_profiles,omitempty"`
|
|
Task string `json:"task"`
|
|
Duration float64 `json:"duration,omitempty"`
|
|
Language string `json:"language,omitempty"`
|
|
NumSpeakers int `json:"num_speakers"`
|
|
Segments []DiarizationSegment `json:"segments"`
|
|
Speakers []DiarizationSpeaker `json:"speakers,omitempty"`
|
|
}
|
|
|
|
// DiarizationResponseFormatType mirrors transcription's response_format
|
|
// pattern: json (default, no per-segment text), verbose_json (adds
|
|
// speakers summary + text when available), and rttm (NIST RTTM rows).
|
|
type DiarizationResponseFormatType string
|
|
|
|
const (
|
|
DiarizationResponseFormatJson DiarizationResponseFormatType = "json"
|
|
DiarizationResponseFormatJsonVerbose DiarizationResponseFormatType = "verbose_json"
|
|
DiarizationResponseFormatRTTM DiarizationResponseFormatType = "rttm"
|
|
)
|