|
1)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89304)
Posted 17 Jul 2018 by Simplex0 Post: I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well. The likelihood of the current WUs running long is not related to the estimated runtime of the BOINC Manager. If the name of the current tasks has the same prefix as those that you had trouble with, that would be a better indicator for you. I have now checked more than 1000 workunits that has finished successfully and only 1 of them took 4 hours while ALL of the invalid aivan workuntis took more that 6 hours to finish. Anyway. It seams that I do not get any more of this kind of workunits so hopefully the problem has already been spotted and taking care of by the staff. |
|
2)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89303)
Posted 16 Jul 2018 by Simplex0 Post: I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well. The likelihood of the current WUs running long is not related to the estimated runtime of the BOINC Manager. If the name of the current tasks has the same prefix as those that you had trouble with, that would be a better indicator for you. The main issue here is not the runtime or credit in this case, it is that a lot of your crunchers resources and YOUR resorses I wasted when a lot of hours of crunching ends up with a result that is invald. |
|
3)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89302)
Posted 16 Jul 2018 by Simplex0 Post: I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well. The work units that was marked as 'Invali' is in my first post in this thread and I had 5 - 6 more of the same later. Why you can't find them I have no idea, ask the staff, maybe they can help you. The lates work units of this kind is here.... http://boinc.bakerlab.org/result.php?resultid=1015234868 http://boinc.bakerlab.org/result.php?resultid=1015234886 http://boinc.bakerlab.org/result.php?resultid=1015234755 http://boinc.bakerlab.org/result.php?resultid=1015234804 They was the first units I run in both my fist and second attempt to crunch a bunch of maybe 100 work units but because the 4 first I run all ended up as 'Invalid' I aborted all the others. I have not received any more of this "avian" workunits lately and I hope I wont. The credit is totally irrelevant in this case, the problem imo is that recourses are wasted when hours of crunching ends up with a result that is Invalid. Luckily I spotted the early and wasted only 40 hours instead of 500 hours. |
|
4)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 89287)
Posted 15 Jul 2018 by Simplex0 Post: In case anyone in the Rosetta staff do care if the members computer runs a lot of workunits for 5 - 6 hours each and that all that work is wasted because the result ends up as "Invalid". The work units all have the word 'aivan' in their name the las time I spotted them I downloaded 200 workuntis but aborted all of the after the 5 first had finished after 5 hours despite my settings for "Target CPU run time" is 2 hours and they all ended up as "Invalid" Here is one of the tasks. Uppgift 1015234886 Namn T1000_full3_aivan_SAVE_ALL_OUT_03_09_677955_4155_0 Arbetsenhet 914799018 Skapades 14 Jul 2018, 5:47:44 UTC Skickad 14 Jul 2018, 6:11:13 UTC Rapporteringstidsgräns 22 Jul 2018, 6:11:13 UTC Mottagit 14 Jul 2018, 16:21:25 UTC Servertillstånd Klar Resultat Valideringsfel Enhetstillstånd Färdig Avsluts status 0 (0x00000000) Dator-ID 3418863 Körtid 6 timmar 11 minsta 41 sekunder CPU-tid 6 timmar 8 minsta 50 sekunder Valideringsstatus Inte godkänd Poäng 0.00 Enhetens störst flyttalshastighet 4.91 GFLOPS Applikations version Rosetta v4.07 windows_intelx86 Peak working set size 235.55 MB Peak swap size 229.32 MB Peak disk usage 514.59 MB Stderr logg <core_client_version>7.10.2</core_client_version> <![CDATA[ <stderr_txt> command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.3.flags -in:file:boinc_wu_zip T1000.3.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 3683496 Starting watchdog... Watchdog active. BOINC:: CPU time: 22129.6s, 14400s + 7200s[2018- 7-14 18:21:12:] :: BOINC WARNING! cannot get file size for default.out.gz: could not open file. Output exists: default.out.gz Size: -1 InternalDecoyCount: 0 (GZ) ----- 0 ----- Stream information inconsistent. Writing W_0000001 ====================================================== DONE :: 1 starting structures 22129.6 cpu seconds This process generated 1 decoys from 1 attempts ====================================================== 18:21:12 (14532): called boinc_finish(0) </stderr_txt> ]]> |
|
5)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89279)
Posted 14 Jul 2018 by Simplex0 Post: Yupp! Same error as always, 4 work units and in total 20 hours of wasted computing, luckily I aborted all the other avian workuntis before they started running and wasted even more recourses. Stderr logg <core_client_version>7.10.2</core_client_version> <![CDATA[ <stderr_txt> command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.3.flags -in:file:boinc_wu_zip T1000.3.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 3683498 Starting watchdog... Watchdog active. BOINC:: CPU time: 22129s, 14400s + 7200s[2018- 7-14 18:20:52:] :: BOINC WARNING! cannot get file size for default.out.gz: could not open file. Output exists: default.out.gz Size: -1 InternalDecoyCount: 0 (GZ) ----- 0 ----- Stream information inconsistent. Writing W_0000001 ====================================================== DONE :: 1 starting structures 22129 cpu seconds This process generated 1 decoys from 1 attempts ====================================================== 18:20:52 (10344): called boinc_finish(0) </stderr_txt> |
|
6)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89277)
Posted 14 Jul 2018 by Simplex0 Post: Ones again I got avian tasks that have been running for 2,5 hours and are estimated to run for 3 - 4 hours more despite that my settings in Rosetta for Target CPU run time is 2 hours. Should I abort them? I have aborted all other avian tasks as my experience is that they are running for a long time an all end up with an error. |
|
7)
Message boards :
Number crunching :
I'm geting lots of errors with Rosetta v4.07
(Message 89262)
Posted 12 Jul 2018 by Simplex0 Post: I think I will abort all v4.07 from now on, this is how the Stderr logg looks Stderr logg <core_client_version>7.10.2</core_client_version> <![CDATA[ <stderr_txt> command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.flags -in:file:boinc_wu_zip T1000.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 1202697 Starting watchdog... Watchdog active. BOINC:: CPU time: 22193.3s, 14400s + 7200s[2018- 7-12 18:11:32:] :: BOINC WARNING! cannot get file size for default.out.gz: could not open file. Output exists: default.out.gz Size: -1 InternalDecoyCount: 0 (GZ) ----- 0 ----- Stream information inconsistent. Writing W_0000001 ====================================================== DONE :: 1 starting structures 22194.4 cpu seconds This process generated 1 decoys from 1 attempts ====================================================== 18:11:32 (10080): called boinc_finish(0) </stderr_txt> <message> upload failure: <file_xfer_error> <file_name>T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5924_0_r2041150700_0</file_name> <error_code>-240 (stat() failed)</error_code> </file_xfer_error> </message> ]]> Seams to be only this type of workunit Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5951_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5959_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5971_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5972_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5991_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4411_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4297_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4122_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4081_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_3741_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_3374_0 Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_2579_0 |
|
8)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 89228)
Posted 6 Jul 2018 by Simplex0 Post: If running fewer concurrent tasks improves credit per minute, it would imply either memory contention, or L2 cache contention. If all of the cores are operating in the same L2 cache, then you can see how that would become the constrained resource. I don't mean to say there is anything wrong with a given computer, just that R@h is very memory intensive. Also others have found that machines with larger L2 caches seem to yield more credit per FLOPS rating per runtime minute. Lets compare Threadripper 1950X, 16 core, L2-cache size 16 x 512 KB and Intel Celeron G1620, 2 core, L2 cache size 512 KB The L2 cache size per core is actually half the size per core on the Intel Celeron G1620 compared with Threadripper 1950X but a Threadripper 1950X, running Rosetta on all cores, is being outperformed by a Intel Celeron G1620. Does that make any sense to you? Take a look at computer ranked as number 51 here https://boinc.bakerlab.org/rosetta/top_hosts.php?sort_by=expavg_credit&offset=40 My experience with Rosetta while running a Threadripper 1950X@4GHz on full load on 31 threads 16 hoursday is as follow..... Target CPU run time 1 hours credit per day = 23400 Target CPU run time 2 hours credit per day = 28100 Target CPU run time 4 hours credit per day = 16700 Target CPU run time 8 hours credit per day = 13700 What credit you will get running a given CPU in Rosetta is a lottery where you have no idea of what you could expect yet you see staffmoderators here claim that the credit given for work done works as it should. |
|
9)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89199)
Posted 1 Jul 2018 by Simplex0 Post: Since many of you seem interested, and mystified by the credit system at R@h, I'll try to clarify a few points. Some other projects simply use CPU seconds, or FLOPS of a task to issue credit. The problem with such a system is that it doesn't reward machines that achieve more work per second due to CPU cache, and memory available. It also depends upon the BOINC Manager to track and report the FLOPS expended on a task, and so some users began to falsify the FLOPS benchmarks and thus the FLOPS of their reported results. If that's the case than you should not run this project using a 16 core, 32 thread Threadripper as it is a waste of money and energy because it is being outperformed by a factor of 2 by a single 2 core Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9]. As you can see here this host https://boinc.bakerlab.org/rosetta/show_host_detail.php?hostid=3391974 is ranked at 45 and is producing slightly more in Rosetta than a 22 core, 44 thread Intel(R) Xeon(R) CPU E5-2696 v4 @ 2.20GHz [Family 6 Model 79 Stepping 1] Thank you for the clarification. |
|
10)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89190)
Posted 29 Jun 2018 by Simplex0 Post: null |
|
11)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89189)
Posted 29 Jun 2018 by Simplex0 Post: I need a working credit system to evaluate which CPU, GPU, OS to use with which project to make it as effective as possible. I have run Folding@home, the GPU application, for years and that seams so work very well regarding credit and how different cards performs on average over time. Now I will start running CPU applications as well to see how that works comparing with this and other projects under BOINC. A major difference with Folding@home as I understand is that their software is made by professional programmers while the various programs in projects under BOINC sometimes are made by programmers that ar far less experience in programming. I remember the difference Cluster Physics made when he rewrote the program running in Milkyway@home and made it run 10000 times faster than the original program. |
|
12)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89184)
Posted 29 Jun 2018 by Simplex0 Post: I need a working credit system to evaluate which CPU, GPU, OS to use with which project to make it as effective as possible. I remember that when I was running SIMAP I could see, based on credit I got, that my AMD FX8350 was more effective than Intel 3770K so based on that I used my FX8350 to run work units under SIMAP. |
|
13)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89181)
Posted 29 Jun 2018 by Simplex0 Post: Here it seams that an Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9] I will try out running cpu work units under Folding@home and try to find other medical projects under BOINC that has a better credit system than this. It is hard to understand that it can be so complicated to make a fair credit system that is based on the work you are produce and nothing else. |
|
14)
Message boards :
Number crunching :
For the betterment of BOINC
(Message 89179)
Posted 29 Jun 2018 by Simplex0 Post: I agree with rcthardcore regarding CreditScrew. Here it seams that an Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9] (2 processorer) ar producing more than twice the credit as I get running an 15 core (32 thread) @4GHz AMD Threadripper on full load or should I take it that a Threadripper is a bad choice regarding efficiency if you want ro run Rosetta? |
©2026 University of Washington
https://www.bakerlab.org