flâneur

openbmb/UltraFeedback · Datasets at Hugging Face

huggingface.co · 3,107 words · saved by 1 readers

UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We then use these prompts to query multiple LLMs (see Table for model lists) and generate 4 different responses for each prompt, resulting in a total of 256k samples. To collect high-quality preference and textual feedback, we design a fine-grained annotation instruction, which contains 4 different aspects, namely instruction-following, truthfulness, honesty and helpfulness. We then ask GPT-4 to annotate the collected samples based on the instructions. We sample 63,967 instructions from 6 public available and high-quality datasets. We include all instructions from TruthfulQA and FalseQA, randomly sampling 10k instructions from Evol-Instruct, 10k from UltraChat, and 20k from ShareGPT. For Flan, we adopt a stratified sampling s

Subset (1) default · 64k rows Split (1) source string classes 9 values instruction string lengths 7 14.5k models sequence completions list correct_answers sequence incorrect_answers sequence evol_instruct Can you write a C++ program that prompts the user to enter the name of a country and checks if it borders the Mediterranean Sea? Here's some starter code to help you out: #include <iostream> #include <string> using namespace std; int main() { string country; // prompt user for input cout << "Enter the name ... [ "alpaca-7b", "pythia-12b", "starchat", "vicuna-33b" ]…

related reading