Papers
AICore idea
To fake a human, stop copying and start fooling the judge
AICore idea
One Transformer Layer Can Match Full RL Training
AICore idea
Teaching image-reading AI to stop and run code when the math gets hard
AICore idea
A small AI watches video like a detective — and beats a model 10x its size
AICore idea
An AI that teaches itself gets sharper but stops exploring
Related concepts
Ideas that show up alongside Reinforcement Learning.
Active PerceptionAdaptive Interleaved ReasoningAgentic Decision-MakingAudio-Visual ReasoningCode GenerationCode-Augmented ReasoningConversational AgentsData Curation And FilteringGraph Path-FindingKnowledge DistillationLanguage Model ReasoningLarge Language Model Post-TrainingLarge Language ModelsLayer-Wise AnalysisLLM-as-a-Judge EvaluationLong Video UnderstandingMathematical ReasoningMultimodal Large Language ModelsMutual InformationNumerical ComputationOmni-Modal AgentOn-Policy Self-DistillationOut-Of-Distribution GeneralizationOutput DiversityParameter-Efficient TrainingPartially Observable Markov Decision ProcessPersonalization SystemsReward ModelingSupervised Fine-TuningTest-Time ScalingTool-Use In AITransformer ArchitectureTuring TestUser Simulation