יש מאמר מעניין של אפל שיצא לפני חודשיים שבוחן בדרכים יצירתיות את היכולות ה"אמיתיות" (שלא מושתתות על שינון) של מודלי השפה בתחום של mathematical reasoning.
להבנתי התוצאות די מאכזבות בגדול (נראה שיש שם הרבה שינון ופחות reasoning).
arxiv.org
להבנתי התוצאות די מאכזבות בגדול (נראה שיש שם הרבה שינון ופחות reasoning).
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Recent advancements in Large Language Models (LLMs) have sparked interest in their formal reasoning capabilities, particularly in mathematics. The GSM8K benchmark is widely used to assess the mathematical reasoning of models on grade-school-level questions. While the performance of LLMs on GSM8K...