Skip to main content
Byser

Data your model can learn from.

Byser designs and delivers purpose-built datasets for AI training and evaluation, with the structure, annotations and metadata your model task requires.

Model task

Built around the model task.

The same raw media can become very different datasets depending on what a model needs to learn. Byser defines the useful unit of data, the relationships between fields and the technical context around each example, so the collection reflects the job it is meant to do.

01

Structured supervision

Transcripts, OCR, temporal segments, labels and other supervisory signals can be attached where they add useful information to the task.

02

Explicit provenance

Source and label provenance can remain visible in the data, helping technical teams distinguish different kinds of supervision instead of treating every field as equivalent.

03

Clear delivery structure

Media, records and metadata can be connected through stable identifiers and machine-readable manifests, reducing the work required to understand how the package fits together.

Dataset collections

Explore our dataset collections

Speech, screen, video, audio, image and multimodal formats, each designed around a distinct modelling problem.

AUDIO01

Natural English Dialogue Sets

Natural multi-speaker conversation with the turn structure, timing and speaker context needed for ASR, diarisation and conversational systems.

AUDIODIALOGUE
View dataset
AUDIO02

Cross-Language Dialogue Sets

Multilingual conversation organised under a consistent data structure while preserving the language-specific information each locale requires.

AUDIOMULTILINGUAL
View dataset
AUDIO03

Accent and Regional Speech Sets

Speech organised across defined accent and regional groups for adaptation, coverage analysis and comparative model evaluation.

AUDIOACCENT COVERAGE
View dataset
AUDIO04

Rare and Low-Resource Language Sets

Speech data for languages and varieties where existing machine-learning resources are limited, fragmented or difficult to standardise.

AUDIOLANGUAGE COVERAGE
View dataset
AUDIO05

Instruction and Command Voice Sessions

Spoken instructions and short follow-up exchanges that preserve intent, correction, confirmation and clarification behaviour.

AUDIOCOMMAND SPEECH
View dataset
AUDIO06

Bilingual Code-Switching Dialogue Sets

Bilingual conversation in which language changes can be represented within turns rather than reduced to a single session label.

AUDIOCODE-SWITCHING
View dataset

Need something more specific?

Start with the model problem. We can define the modality, unit of observation, annotations, metadata and delivery structure around the way the data will actually be used.

Explore all collections

Technical delivery

A dataset should explain itself.

Depending on the collection, delivery can include task-specific records, machine-readable manifests, provenance fields, stable identifiers, SHA-256 hashes, MLCommons Croissant 1.0 metadata and loading instructions. The point is not to add more files. It is to make the relationship between the source data and the model task clear.

How it works

Tell us what your model needs.

Share the task, the data conditions that matter and the structure you need to work with. We will use that to define the collection.