Custom dataset service · On request
Custom web datasets,built to your brief
Commission custom web datasets around the sources, fields and decisions that matter to your team. QuanticData works with you to agree a collection brief, review a representative sample and define quality checks, provenance and CSV or JSON delivery. Optional refresh is scoped around your workflow.
Start with your sources and desired output.
By Aldo Morese · Updated
{
"property_ref": "string",
"source_url": "URL",
"observed_at": "ISO 8601 timestamp",
"check_in": "date",
"check_out": "date",
"occupancy": "agreed guest/room structure",
"room_type": "string",
"rate_plan": "string",
"currency": "currency code",
"stay_total": "decimal or null",
"tax_basis": "string",
"availability": "string or null"
}
A starting point for discussion. Final fields and source coverage are agreed for your project.
Define useful records
Sources and schema, agreed upfront.
Bring representative URLs, markets, languages and the question your dataset should answer. We discuss source feasibility and agree the collection boundaries. Define each field, its type, the unit of a record and which values are essential. Separate a company, a property and an individual offer before combining their data.
For hotel offers, agree stay dates, occupancy, room and rate plans, cancellation terms, currency and tax treatment. These details make comparisons meaningful. For other briefs, explore our specialist workflows for company data collection and market research data.
Review before expanding
A sample makes the brief testable.
Agree sample work, review responsibilities and acceptance criteria before the broader collection. Sampling should include source variations that could change the output.
- Check the fields.Review required values, types and formatting against your intended use. Record gaps and ambiguities for discussion.
- Check the records.Agree duplicate handling and matching rules. Flag uncertain hotel or product matches instead of merging them silently.
- Keep the evidence.Include source URLs and observation times in the agreed schema. Distinguish collection failures, unavailable offers and genuinely absent fields.
Fit your receiving workflow
CSV or JSON, with refresh by agreement.
Choose a flat CSV structure for tabular analysis or an agreed JSON structure for application use. Define field names, missing-value conventions, delivery destination and accompanying notes. Provenance helps your team trace an observation to its source and collection time.
Request a snapshot or discuss recurring collection. Refresh cadence, source changes and delivery handling are agreed separately. If you need comparison rules and change notifications, explore custom price monitoring and webhooks.
Choose the workflow that fits your brief.
This commissioned service is available on request. The separate Quantic AI prompt-to-dataset software is in development. For an existing extraction format your team can integrate directly, browse ready-made collectors.
For AI or retrieval projects, define the records and provenance your workflow needs. Collecting those records is distinct from annotation, labeling or model evaluation; any additional work needs a separate discussion.
Before you send a brief
Custom datasets FAQ
What is a custom web dataset?
A custom web dataset is a collection of records gathered from agreed web sources to answer your particular brief. You specify the fields, coverage and intended use. We scope the collection and delivery with you, rather than asking you to filter a fixed catalog of existing downloads.
What should I include in my dataset request?
Send representative source URLs, the question you want to answer and a draft list of fields. Include markets, languages, date ranges and your preferred delivery format. An example of the table your team wants to use helps us discuss feasibility, sampling and the work involved.
Can we review a sample before the full collection?
Sample work and review are agreed as part of the project scope. A representative sample helps both teams check field definitions, coverage and difficult source variations. Agree acceptance criteria and how corrections will be handled before deciding whether to extend collection to the remaining sources.
How do you handle missing fields and duplicates?
We agree required fields, validation rules and record identifiers before collection. The scope defines how to represent missing values, flag uncertain matches and identify duplicates. Unavailable data stays distinguishable from an observed value; a missing hotel price, for example, should not become a zero price.
Can you deliver CSV or JSON and refresh the dataset?
CSV or JSON delivery can be agreed around your schema and receiving workflow. Discuss the destination, file structure and provenance fields in the brief. Recurring refresh is optional: its cadence, source coverage, change handling and delivery arrangements need their own agreed scope rather than automatic activation.
Is this the Quantic AI prompt-to-dataset product?
No. This is a commissioned service available on request, with a brief, scope discussion and agreed deliverables. Quantic AI is separate prompt-to-dataset software in development. Contacting us about a custom dataset starts a project discussion; it does not activate that software or create an instant download.
How is a custom dataset priced?
The quote depends on the agreed sources, extraction work, schema, quality checks and delivery requirements. Optional refresh and additional integrations also affect scope. Send your brief so we can discuss feasibility and commercial terms before work starts. Standard API rates are not a quote for a commissioned dataset.
Tell us what the dataset should answer.
Email your source examples, fields and delivery preferences. We will discuss a scope and quote for your project.