Papers
AICore idea
To fake a human, stop copying and start fooling the judge
AICore idea
Teaching robots by thumbs-up: a smarter way to pick what to try next
AIMentioned
An AI that teaches itself to see and draw — with no human labels
Related concepts
Ideas that show up alongside Reward Modeling.
Autoregressive ModelsConversational AgentsDiffusion ModelsEnsemble MethodsEpistemic Uncertainty EstimationExplorationImage GenerationLarge Language ModelsLLM-as-a-Judge EvaluationModel-Based Reinforcement LearningPersonalization SystemsPreference-Based Reinforcement LearningRegret BoundsReinforcement LearningSample EfficiencySelf-Consistency RewardsSelf-Evolving LearningTuring TestUnified Multimodal Understanding And GenerationUser SimulationVisual Question Answering