Papers
IAIdea central
To fake a human, stop copying and start fooling the judge
IAIdea central
Entrenar una sola capa iguala a entrenar todo el modelo
IAIdea central
Teaching image-reading AI to stop and run code when the math gets hard
IAIdea central
A small AI watches video like a detective — and beats a model 10x its size
IAIdea central
An AI that teaches itself gets sharper but stops exploring
Conceptos relacionados
Ideas que aparecen junto a Reinforcement Learning.
Active PerceptionAdaptive Interleaved ReasoningAgentic Decision-MakingAudio-Visual ReasoningCode GenerationCode-Augmented ReasoningConversational AgentsData Curation And FilteringGraph Path-FindingKnowledge DistillationLanguage Model ReasoningLarge Language Model Post-TrainingLarge Language ModelsLayer-Wise AnalysisLLM-as-a-Judge EvaluationLong Video UnderstandingMathematical ReasoningMultimodal Large Language ModelsMutual InformationNumerical ComputationOmni-Modal AgentOn-Policy Self-DistillationOut-Of-Distribution GeneralizationOutput DiversityParameter-Efficient TrainingPartially Observable Markov Decision ProcessPersonalization SystemsReward ModelingSupervised Fine-TuningTest-Time ScalingTool-Use In AITransformer ArchitectureTuring TestUser Simulation