Posts by Simplex0

1) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89304)
Posted 17 Jul 2018 by Simplex0
Post:
I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well. The likelihood of the current WUs running long is not related to the estimated runtime of the BOINC Manager. If the name of the current tasks has the same prefix as those that you had trouble with, that would be a better indicator for you.


I have now checked more than 1000 workunits that has finished successfully and only 1 of them took 4 hours while ALL of the invalid aivan workuntis took more that 6 hours to finish.
Anyway. It seams that I do not get any more of this kind of workunits so hopefully the problem has already been spotted and taking care of by the staff.
2) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89303)
Posted 16 Jul 2018 by Simplex0
Post:
I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well. The likelihood of the current WUs running long is not related to the estimated runtime of the BOINC Manager. If the name of the current tasks has the same prefix as those that you had trouble with, that would be a better indicator for you.


The main issue here is not the runtime or credit in this case, it is that a lot of your crunchers resources and YOUR resorses I wasted when a lot of hours of crunching ends up with a result that is invald.
3) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89302)
Posted 16 Jul 2018 by Simplex0
Post:
I have no details on the specific WUs or issues they are having. But I wanted everyone to know that BOINC Manager's "estimated runtime" is really based on history, not the present. So, regardless of the name or likely success of current WUs, if BOINC Manager has a recent history with WUs taking 3 or 4 hours longer than the runtime preference, it will "estimate" future WUs will take 3 to 4 hours longer as well.


For me the problem is not the runtime of wus (i know the decoy's question), but the validation error.


I found 1 "Invalid" run that you were granted 587.11 credits. Were there others that were a problem for you?
Seems like the 587 credits were similar to the other valid jobs.



name T1000_full3_aivan_SAVE_ALL_OUT_03_09_677955_5874
application Rosetta
created 14 Jul 2018, 8:17:20 UTC
canonical result 1015258491
granted credit 587.11

https://boinc.bakerlab.org/workunit.php?wuid=914821393


The work units that was marked as 'Invali' is in my first post in this thread and I had 5 - 6 more of the same later.
Why you can't find them I have no idea, ask the staff, maybe they can help you.

The lates work units of this kind is here....

http://boinc.bakerlab.org/result.php?resultid=1015234868
http://boinc.bakerlab.org/result.php?resultid=1015234886
http://boinc.bakerlab.org/result.php?resultid=1015234755
http://boinc.bakerlab.org/result.php?resultid=1015234804

They was the first units I run in both my fist and second attempt to crunch a bunch of maybe 100 work units but because the 4 first I run all ended up as 'Invalid' I aborted all the others.
I have not received any more of this "avian" workunits lately and I hope I wont.
The credit is totally irrelevant in this case, the problem imo is that recourses are wasted when hours of crunching ends up with a result that is Invalid.
Luckily I spotted the early and wasted only 40 hours instead of 500 hours.
4) Message boards : Number crunching : Problems and Technical Issues with Rosetta@home (Message 89287)
Posted 15 Jul 2018 by Simplex0
Post:
In case anyone in the Rosetta staff do care if the members computer runs a lot of workunits for 5 - 6 hours each and that all that work is wasted because the result ends up as "Invalid".

The work units all have the word 'aivan' in their name the las time I spotted them I downloaded 200 workuntis but aborted all of the after the 5 first had finished after 5 hours despite my settings for "Target CPU run time" is 2 hours and they all ended up as "Invalid"

Here is one of the tasks.


Uppgift 1015234886

Namn T1000_full3_aivan_SAVE_ALL_OUT_03_09_677955_4155_0
Arbetsenhet 914799018
Skapades 14 Jul 2018, 5:47:44 UTC
Skickad 14 Jul 2018, 6:11:13 UTC
Rapporteringstidsgräns 22 Jul 2018, 6:11:13 UTC
Mottagit 14 Jul 2018, 16:21:25 UTC
Servertillstånd Klar
Resultat Valideringsfel
Enhetstillstånd Färdig
Avsluts status 0 (0x00000000)
Dator-ID 3418863
Körtid 6 timmar 11 minsta 41 sekunder
CPU-tid 6 timmar 8 minsta 50 sekunder
Valideringsstatus Inte godkänd
Poäng 0.00
Enhetens störst flyttalshastighet 4.91 GFLOPS
Applikations version Rosetta v4.07
windows_intelx86
Peak working set size 235.55 MB
Peak swap size 229.32 MB
Peak disk usage 514.59 MB

Stderr logg
<core_client_version>7.10.2</core_client_version>
<![CDATA[
<stderr_txt>
command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.3.flags -in:file:boinc_wu_zip T1000.3.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 3683496
Starting watchdog...
Watchdog active.
BOINC:: CPU time: 22129.6s, 14400s + 7200s[2018- 7-14 18:21:12:] :: BOINC
WARNING! cannot get file size for default.out.gz: could not open file.
Output exists: default.out.gz Size: -1
InternalDecoyCount: 0 (GZ)
-----
0
-----
Stream information inconsistent.
Writing W_0000001
======================================================
DONE :: 1 starting structures 22129.6 cpu seconds
This process generated 1 decoys from 1 attempts
======================================================
18:21:12 (14532): called boinc_finish(0)

</stderr_txt>
]]>
5) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89279)
Posted 14 Jul 2018 by Simplex0
Post:
Yupp!
Same error as always, 4 work units and in total 20 hours of wasted computing, luckily I aborted all the other avian workuntis before they started running and wasted even more recourses.

Stderr logg
<core_client_version>7.10.2</core_client_version>
<![CDATA[
<stderr_txt>
command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.3.flags -in:file:boinc_wu_zip T1000.3.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 3683498
Starting watchdog...
Watchdog active.
BOINC:: CPU time: 22129s, 14400s + 7200s[2018- 7-14 18:20:52:] :: BOINC
WARNING! cannot get file size for default.out.gz: could not open file.
Output exists: default.out.gz Size: -1
InternalDecoyCount: 0 (GZ)
-----
0
-----
Stream information inconsistent.
Writing W_0000001
======================================================
DONE :: 1 starting structures 22129 cpu seconds
This process generated 1 decoys from 1 attempts
======================================================
18:20:52 (10344): called boinc_finish(0)

</stderr_txt>
6) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89277)
Posted 14 Jul 2018 by Simplex0
Post:
Ones again I got avian tasks that have been running for 2,5 hours and are estimated to run for 3 - 4 hours more despite that my settings in Rosetta for Target CPU run time is 2 hours.
Should I abort them?
I have aborted all other avian tasks as my experience is that they are running for a long time an all end up with an error.
7) Message boards : Number crunching : I'm geting lots of errors with Rosetta v4.07 (Message 89262)
Posted 12 Jul 2018 by Simplex0
Post:
I think I will abort all v4.07 from now on, this is how the Stderr logg looks

Stderr logg
<core_client_version>7.10.2</core_client_version>
<![CDATA[
<stderr_txt>
command: projects/boinc.bakerlab.org_rosetta/rosetta_4.07_windows_intelx86.exe @T1000.flags -in:file:boinc_wu_zip T1000.zip -nstruct 10000 -cpu_run_time 28800 -watchdog -boinc:max_nstruct 600 -checkpoint_interval 120 -mute all -database minirosetta_database -in::file::zip minirosetta_database.zip -boinc::watchdog -run::rng mt19937 -constant_seed -jran 1202697
Starting watchdog...
Watchdog active.
BOINC:: CPU time: 22193.3s, 14400s + 7200s[2018- 7-12 18:11:32:] :: BOINC
WARNING! cannot get file size for default.out.gz: could not open file.
Output exists: default.out.gz Size: -1
InternalDecoyCount: 0 (GZ)
-----
0
-----
Stream information inconsistent.
Writing W_0000001
======================================================
DONE :: 1 starting structures 22194.4 cpu seconds
This process generated 1 decoys from 1 attempts
======================================================
18:11:32 (10080): called boinc_finish(0)

</stderr_txt>
<message>
upload failure: <file_xfer_error>
<file_name>T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5924_0_r2041150700_0</file_name>
<error_code>-240 (stat() failed)</error_code>
</file_xfer_error>
</message>
]]>

Seams to be only this type of workunit

Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5951_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5959_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5971_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5972_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_5991_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4411_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4297_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4122_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_4081_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_3741_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_3374_0
Namn T1000_full_aivan_SAVE_ALL_OUT_03_09_677708_2579_0
8) Message boards : Number crunching : Problems and Technical Issues with Rosetta@home (Message 89228)
Posted 6 Jul 2018 by Simplex0
Post:
If running fewer concurrent tasks improves credit per minute, it would imply either memory contention, or L2 cache contention. If all of the cores are operating in the same L2 cache, then you can see how that would become the constrained resource. I don't mean to say there is anything wrong with a given computer, just that R@h is very memory intensive. Also others have found that machines with larger L2 caches seem to yield more credit per FLOPS rating per runtime minute.


Lets compare

Threadripper 1950X, 16 core, L2-cache size 16 x 512 KB

and

Intel Celeron G1620, 2 core, L2 cache size 512 KB

The L2 cache size per core is actually half the size per core on the Intel Celeron G1620 compared with Threadripper 1950X but a Threadripper 1950X, running Rosetta on all cores, is being outperformed by a Intel Celeron G1620.
Does that make any sense to you?

Take a look at computer ranked as number 51 here https://boinc.bakerlab.org/rosetta/top_hosts.php?sort_by=expavg_credit&offset=40

My experience with Rosetta while running a Threadripper 1950X@4GHz on full load on 31 threads 16 hoursday is as follow.....

Target CPU run time 1 hours credit per day = 23400
Target CPU run time 2 hours credit per day = 28100
Target CPU run time 4 hours credit per day = 16700
Target CPU run time 8 hours credit per day = 13700

What credit you will get running a given CPU in Rosetta is a lottery where you have no idea of what you could expect yet you see staffmoderators here claim that the credit given for work done works as it should.
9) Message boards : Number crunching : For the betterment of BOINC (Message 89199)
Posted 1 Jul 2018 by Simplex0
Post:
Since many of you seem interested, and mystified by the credit system at R@h, I'll try to clarify a few points. Some other projects simply use CPU seconds, or FLOPS of a task to issue credit. The problem with such a system is that it doesn't reward machines that achieve more work per second due to CPU cache, and memory available. It also depends upon the BOINC Manager to track and report the FLOPS expended on a task, and so some users began to falsify the FLOPS benchmarks and thus the FLOPS of their reported results.

To avoid these problems and design a credit system that would work going forward through new generations of hardware and capabilities, a system was adopted whereby credit granted is based on what OTHER user's machines have claimed for how difficult it is to complete each decoy (or model) of a task. The average of the results reported before you is used, rather than an arbitrary benchmark. So the credit granted does not favor a specific CPU type or operating system. It truly is based on completed work. Some tasks use different algorithms to compute their results. Assessing the effectiveness of a variety of algorithms is how the science progresses. This results in variations on computing resources required to complete them. Each type of task self-adjusts the credit it awards.

This makes a credit system that is great at reflecting the value of work to the Project Team, but makes it more difficult to benchmark your own machines. To do that well, you would have to run the same task, from the same starting seed for each benchmark run, and then compare CPU and wall-clock time, as well as the number of decoys produced. The BOINC Manager doesn't have a means of easily supporting this type of usage. Running tasks for other projects or a different class of proteins only tells you about those things. One machine might be faster than another at one type of work, and not with another type. The R@h credit system reflects the value of the outcome, without bias.


If that's the case than you should not run this project using a 16 core, 32 thread Threadripper as it is a waste of money and energy because it is being outperformed by a factor of 2 by a single 2 core Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9].

As you can see here this host https://boinc.bakerlab.org/rosetta/show_host_detail.php?hostid=3391974 is ranked at 45 and is producing slightly more in Rosetta than a 22 core, 44 thread Intel(R) Xeon(R) CPU E5-2696 v4 @ 2.20GHz [Family 6 Model 79 Stepping 1]


Thank you for the clarification.
10) Message boards : Number crunching : For the betterment of BOINC (Message 89190)
Posted 29 Jun 2018 by Simplex0
Post:
null
11) Message boards : Number crunching : For the betterment of BOINC (Message 89189)
Posted 29 Jun 2018 by Simplex0
Post:
I need a working credit system to evaluate which CPU, GPU, OS to use with which project to make it as effective as possible.
I remember that when I was running SIMAP I could see, based on credit I got, that my AMD FX8350 was more effective than Intel 3770K so based on that I used my FX8350 to run work units under SIMAP.


I think everyone would like that. 8-) Even those running the projects. It is very difficult to do.


On Win10 I use the Windows TASK MANAGER to monitor especially the DISK and NET usage. I check out any activity on either one of those.
I use BoincTasks to monitor BOINC WU. I watch the CPU usage of GPU task and MEMORY use of WU.
Einstein GPU tasks take 99% of a CPU and about 1GB of virtual memory on my machine. MilkyWay uses 0% CPU.
Rosetta WU MEMORY usage seems to bloat up to 1GB+ just before CHECKPOINTING.

I use GPU-Z to monitor the loads on the GPU and memory traffic.

I back off the CPU available to CPU jobs until I see the TASK MANAGER CPU drop below 100% and then add a CPU back in. For my 12-core i9-7920X, I had to exclude 2 threads from BOINC CPU WU.

The PROJECTS tend to optimize for a single thread which likely negatively affects the total throughput.
PrimeGrid aggressively prefetches all the data into the CPU caches before starting to crunch.
If you run 1 PrimeGrid WU per CORE, you get shortest compute time.
If you run 1 PrimeGrid WU per THREAD, there is an 80% slow down. You do more work if you use all the THREADS. Just slower.

Many projects break off a fixed length problem and allocate to the volunteer.

Rosetta gives a computer a LONG WU problem and a TIME PERIOD to work on it. If you set your computer TIME to 24 hours, you will see many WU finish much sooner than 24 hours. When I was looking at Rosetta performance, I used a shorter fixed problem where the run time was repeatable and comparable.

Rosetta has a big code footprint. They compile with high optimizations and aggressive INLINE options.
Rosetta does very little floating point. When it does, it does 3D coordinate floating point LOAD, ADD/SUB/MUL, STORE .... 3 times in a row.

When I looked at the new 4.0 version, It spends a significant amount of time spinning on LOCKS. 4.0 seems to do even less floating point.


I have run Folding@home, the GPU application, for years and that seams so work very well regarding credit and how different cards performs on average over time.
Now I will start running CPU applications as well to see how that works comparing with this and other projects under BOINC.

A major difference with Folding@home as I understand is that their software is made by professional programmers while the various programs in projects under BOINC
sometimes are made by programmers that ar far less experience in programming.
I remember the difference Cluster Physics made when he rewrote the program running in Milkyway@home and made it run 10000 times faster than the original program.
12) Message boards : Number crunching : For the betterment of BOINC (Message 89184)
Posted 29 Jun 2018 by Simplex0
Post:
I need a working credit system to evaluate which CPU, GPU, OS to use with which project to make it as effective as possible.
I remember that when I was running SIMAP I could see, based on credit I got, that my AMD FX8350 was more effective than Intel 3770K so based on that I used my FX8350 to run work units under SIMAP.
13) Message boards : Number crunching : For the betterment of BOINC (Message 89181)
Posted 29 Jun 2018 by Simplex0
Post:
Here it seams that an Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9]
(2 processorer) ar producing more than twice the credit as I get running an 15 core (32 thread) @4GHz AMD Threadripper on full load or should I take it that
a Threadripper is a bad choice regarding efficiency if you want ro run Rosetta?

That is my problem, as I use credits to check my equipment. I never know whether what I am seeing is real, or a figment of the imagination.



I will try out running cpu work units under Folding@home and try to find other medical projects under BOINC that has a better credit system than this.
It is hard to understand that it can be so complicated to make a fair credit system that is based on the work you are produce and nothing else.
14) Message boards : Number crunching : For the betterment of BOINC (Message 89179)
Posted 29 Jun 2018 by Simplex0
Post:
I agree with rcthardcore regarding CreditScrew.
Here it seams that an Intel(R) Celeron(R) CPU G1620 @ 2.70GHz [Family 6 Model 58 Stepping 9]
(2 processorer) ar producing more than twice the credit as I get running an 15 core (32 thread) @4GHz AMD Threadripper on full load or should I take it that
a Threadripper is a bad choice regarding efficiency if you want ro run Rosetta?






©2026 University of Washington
https://www.bakerlab.org