Статьи
ИИКлючевая идея
To fake a human, stop copying and start fooling the judge
ИИКлючевая идея
Обучите один слой нейросети — и получите почти всю выгоду от полного обучения
ИИКлючевая идея
ИИ научили писать код прямо посреди размышлений — и точность выросла на 6 пунктов
ИИКлючевая идея
A small AI watches video like a detective — and beats a model 10x its size
ИИКлючевая идея
Модель, обучаясь у самой себя, теряет разнообразие решений
Смежные понятия
Идеи, которые идут в паре с понятием «Reinforcement Learning».
Active PerceptionAdaptive Interleaved ReasoningAgentic Decision-MakingAudio-Visual ReasoningCode GenerationCode-Augmented ReasoningConversational AgentsData Curation And FilteringGraph Path-FindingKnowledge DistillationLanguage Model ReasoningLarge Language Model Post-TrainingLarge Language ModelsLayer-Wise AnalysisLLM-as-a-Judge EvaluationLong Video UnderstandingMathematical ReasoningMultimodal Large Language ModelsMutual InformationNumerical ComputationOmni-Modal AgentOn-Policy Self-DistillationOut-Of-Distribution GeneralizationOutput DiversityParameter-Efficient TrainingPartially Observable Markov Decision ProcessPersonalization SystemsReward ModelingSupervised Fine-TuningTest-Time ScalingTool-Use In AITransformer ArchitectureTuring TestUser Simulation