SparkMoE - מערכת להרצת מודלי AI גדולים ועוצמתיים באופן מקומי על חומרה חלשה, עדכון - גרסה V0.3.0!
-
@CSS-0 זה פחות חדיש מהNVME אבל גם מגיע למהירויות מטורפות, אולי לא מספיק לAI? אין לי מושג (יצא לי להעביר מNVME לSATA במהירות 700 מ"ב בשניה)
@החכם-השלם מה זה 'מספיק'?
גם SSD אמור לעבוד (אלא אם המנגנון העברת נתונים עצמו שונה, אני לא כל כך מבין בזה..) רק שזה יהיה איטי להחריד, כאן זה שונה כי זה קורס..
מסתמא זה בCPU או ראם, אבל מצד שני מודלים מקומיים רגילים בlama.cpp אתה טוען שאתה כן מצליח להריץ.. -
@החכם-השלם מה זה 'מספיק'?
גם SSD אמור לעבוד (אלא אם המנגנון העברת נתונים עצמו שונה, אני לא כל כך מבין בזה..) רק שזה יהיה איטי להחריד, כאן זה שונה כי זה קורס..
מסתמא זה בCPU או ראם, אבל מצד שני מודלים מקומיים רגילים בlama.cpp אתה טוען שאתה כן מצליח להריץ..@המלאך זה הלוגים, ותמיד אני מריץ עליו מודלים (אני מפעיל עליו מערכת טלפונית של AI מבוססת על מודל של 8B וזה עובד תקין וחלק בלי להאט לי את המחשב מעולם, המהירות אינטרנט זה 90MB לשניה, והמהירות שזה מחזיר למשתמש את התשובה זה ממוצע של 7 שניות, (וזה כולל המרה של הטקסט שהמודל מחזיר לשמע עם BlueTTS) אז אני לא חושב שאמור להיות בו בעיה...
בכל מקרה זה הלוגים:0.00.009.579 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.023.223 I srv load_model: loading model 'C:\Users\USER\Downloads\gemma-4-26B-A4B-it-UD-IQ2_M.gguf' 0.01.271.513 W common_fit_params: failed to fit params to free device memory: was unable to fit model into system memory by reducing context, abort 0.02.162.598 W load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.165.271 W load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.177.253 W load: control-looking token: 1 '<eos>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.215.972 W load: special_eog_ids contains '<|tool_response>', removing '</s>' token from EOG list 1.02.461.952 W cmn common_init_: KV cache shifting is not supported for this context, disabling KV cache shifting -
@המלאך זה הלוגים, ותמיד אני מריץ עליו מודלים (אני מפעיל עליו מערכת טלפונית של AI מבוססת על מודל של 8B וזה עובד תקין וחלק בלי להאט לי את המחשב מעולם, המהירות אינטרנט זה 90MB לשניה, והמהירות שזה מחזיר למשתמש את התשובה זה ממוצע של 7 שניות, (וזה כולל המרה של הטקסט שהמודל מחזיר לשמע עם BlueTTS) אז אני לא חושב שאמור להיות בו בעיה...
בכל מקרה זה הלוגים:0.00.009.579 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.023.223 I srv load_model: loading model 'C:\Users\USER\Downloads\gemma-4-26B-A4B-it-UD-IQ2_M.gguf' 0.01.271.513 W common_fit_params: failed to fit params to free device memory: was unable to fit model into system memory by reducing context, abort 0.02.162.598 W load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.165.271 W load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.177.253 W load: control-looking token: 1 '<eos>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.215.972 W load: special_eog_ids contains '<|tool_response>', removing '</s>' token from EOG list 1.02.461.952 W cmn common_init_: KV cache shifting is not supported for this context, disabling KV cache shifting -
@המלאך זה הלוגים, ותמיד אני מריץ עליו מודלים (אני מפעיל עליו מערכת טלפונית של AI מבוססת על מודל של 8B וזה עובד תקין וחלק בלי להאט לי את המחשב מעולם, המהירות אינטרנט זה 90MB לשניה, והמהירות שזה מחזיר למשתמש את התשובה זה ממוצע של 7 שניות, (וזה כולל המרה של הטקסט שהמודל מחזיר לשמע עם BlueTTS) אז אני לא חושב שאמור להיות בו בעיה...
בכל מקרה זה הלוגים:0.00.009.579 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.023.223 I srv load_model: loading model 'C:\Users\USER\Downloads\gemma-4-26B-A4B-it-UD-IQ2_M.gguf' 0.01.271.513 W common_fit_params: failed to fit params to free device memory: was unable to fit model into system memory by reducing context, abort 0.02.162.598 W load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.165.271 W load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.177.253 W load: control-looking token: 1 '<eos>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.215.972 W load: special_eog_ids contains '<|tool_response>', removing '</s>' token from EOG list 1.02.461.952 W cmn common_init_: KV cache shifting is not supported for this context, disabling KV cache shifting -
@המלאך זה הלוגים, ותמיד אני מריץ עליו מודלים (אני מפעיל עליו מערכת טלפונית של AI מבוססת על מודל של 8B וזה עובד תקין וחלק בלי להאט לי את המחשב מעולם, המהירות אינטרנט זה 90MB לשניה, והמהירות שזה מחזיר למשתמש את התשובה זה ממוצע של 7 שניות, (וזה כולל המרה של הטקסט שהמודל מחזיר לשמע עם BlueTTS) אז אני לא חושב שאמור להיות בו בעיה...
בכל מקרה זה הלוגים:0.00.009.579 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.023.223 I srv load_model: loading model 'C:\Users\USER\Downloads\gemma-4-26B-A4B-it-UD-IQ2_M.gguf' 0.01.271.513 W common_fit_params: failed to fit params to free device memory: was unable to fit model into system memory by reducing context, abort 0.02.162.598 W load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.165.271 W load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.177.253 W load: control-looking token: 1 '<eos>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.215.972 W load: special_eog_ids contains '<|tool_response>', removing '</s>' token from EOG list 1.02.461.952 W cmn common_init_: KV cache shifting is not supported for this context, disabling KV cache shifting -
@CSS-0 כנראה אין לך מספיק ram פנוי בשביל בפרמטרים הפעילים. תנסה לשחק עם גודל המטמון בממשק הבקרה.
-
-
@CSS-0

אולי באמת... -
-
@המלאך זה הלוגים, ותמיד אני מריץ עליו מודלים (אני מפעיל עליו מערכת טלפונית של AI מבוססת על מודל של 8B וזה עובד תקין וחלק בלי להאט לי את המחשב מעולם, המהירות אינטרנט זה 90MB לשניה, והמהירות שזה מחזיר למשתמש את התשובה זה ממוצע של 7 שניות, (וזה כולל המרה של הטקסט שהמודל מחזיר לשמע עם BlueTTS) אז אני לא חושב שאמור להיות בו בעיה...
בכל מקרה זה הלוגים:0.00.009.579 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) 0.00.023.223 I srv load_model: loading model 'C:\Users\USER\Downloads\gemma-4-26B-A4B-it-UD-IQ2_M.gguf' 0.01.271.513 W common_fit_params: failed to fit params to free device memory: was unable to fit model into system memory by reducing context, abort 0.02.162.598 W load: control-looking token: 212 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.165.271 W load: control-looking token: 50 '<|tool_response>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.177.253 W load: control-looking token: 1 '<eos>' was not control-type; this is probably a bug in the model. its type will be overridden 0.02.215.972 W load: special_eog_ids contains '<|tool_response>', removing '</s>' token from EOG list 1.02.461.952 W cmn common_init_: KV cache shifting is not supported for this context, disabling KV cache shifting -
@פתרון-AI הכי פשוט זה lm סטודיו.
יש מדריך של @חובבן-מקצועי
חפש באחד הנושאים שלו.עריכה: https://mitmachim.top/topic/95704/מדריך-מדריך-בסיסי-להתקנה-ושימוש-במודלים-ב-lm-studio
-
@פתרון-AI עזוב זה משהו שאני מפיץ במעגל סגור יש ממש אנשים בודדים שנתתי בפרטי, חוץ מזה השיעור שלי בישיבה מחובר לזה וטו לא..
שלום! נראה שהשיחה הזו מעניינת אותך, אבל עדיין אין לך חשבון.
נמאס לכם לגלול בין אותם הפוסטים בכל ביקור? כשנרשמים לחשבון, תמיד תחזרו בדיוק למקום שבו הייתם קודם, ותוכלו לבחור לקבל התראות על תגובות חדשות (בין אם במייל, ובין אם בהתראת פוש). תוכלו גם לשמור סימניות ולפרגן ב-upvote לפוסטים כדי להביע הערכה לחברי קהילה אחרים.
בעזרת התרומה שלך, הפוסט הזה יכול להיות אפילו טוב יותר 💗
הרשמה התחברות