|
1)
Message boards :
Number crunching :
No new work
(Message 64624)
Posted 29 Dec 2009 by netwraith
Post: Server Status shows lots of stuff down .... The front page says they are working on the SAN, but, this is ridiculous if they are still in the process.. They should, at least, post an update... And if it's not the SAN work, they are not paying too much attention over the holiday week.. Perhaps a case of food coma.... |
|
2)
Message boards :
Number crunching :
Minirosetta v1.45 bug thread
(Message 57732)
Posted 9 Dec 2008 by netwraith
Post: I don't know about these units passing on Intel CORE cpu's and more memory... I am having dozens of these cs_vanilla units bomb out on machines with dual CORE2 quad XEONS and 16GB of RAM... I think these machines are big enough to handle anything out there. And I was running 14 of them up to a day or two ago... Most are still crunching, but, are starting to wind down so that they can be part of a compute farm... So, I think these is something wrong in these units or the v1.45 of mini.... |
|
3)
Message boards :
Number crunching :
Problems with web site
(Message 57508)
Posted 2 Dec 2008 by netwraith
Post: For those that are still having issues with your boinc client not getting the new scheduler url from our master url and if you would rather not detach and reattach the project which should fix the issue if updating doesn't. You can do the following: Have your DNS guy create a CNAME alias for schedular.bakaerlab.org and point that to the A record of whatever machine will be doing the scheduling. i.e. $origin bakerlab.org. schedular CNAME srv4.bakerlab.org. (don't forget the trailing period) Then publish the permanent URL for the schedular as http://schedular.bakerlab.org/rosetta_cgi/cgi Then if you ever need to change the machine that runs the schedular, it's only a global DNS change and does not impact any client... If you set the refresh and timeouts short on the zone these changes can be made to propagate really quickly too... sometimes in a couple of minutes.. |
|
4)
Message boards :
Number crunching :
Problems with web site
(Message 57449)
Posted 2 Dec 2008 by netwraith
Post: Mod.Sense suggested a quick fix for this by using the Redirect directive on our webserver. The main question now is does the client handle redirects so please let me know if this fixes the issue or not. The update button did not help, but, programming in a non existent proxy connection and hitting the 'retry communications / do communications' button caused the client to think it was disconnected. After repairing the proxy config to it's prior setting and allowing the deferal clock to timeout (60 seconds in this case), the first thing the client did was to download the scheduler list. This did repair the problem for this client... Now the only problem is that the acknowledged work-units all have 'pending' credits... I have not seen that in a while here..... |
|
5)
Message boards :
Number crunching :
Problems with web site
(Message 57436)
Posted 2 Dec 2008 by netwraith
Post: Now getting a message from BOINC Mon 01 Dec 2008 07:49:39 PM EST|rosetta@home|Message from server: Server error: can't attach shared memory This has been happening the last couple of hours... |
|
6)
Message boards :
Number crunching :
Minirosetta v1.40 bug thread
(Message 57100)
Posted 20 Nov 2008 by netwraith
Post: I have several of these loopbuild_minimalist_core3_homo_bench- .... tasks and several of them are way overtime... get to just under 10 minutes to go and stay that way for hours.... What could be up with these ??? All of my machines are Linux 2.6 kernels... Fedora/RedHat EL/CentOS |
|
7)
Message boards :
Number crunching :
trouble getting new work units
(Message 55162)
Posted 18 Aug 2008 by netwraith
Post: -- <Edit> Just started getting units... never mind... |
|
8)
Message boards :
Number crunching :
Minirosetta v1.32 bug thread
(Message 54988)
Posted 7 Aug 2008 by netwraith
Post: Task reported too late to validate ????? found the problem for this one.... copied directory from another cruncher and forgot to dump the client_state.... my bad... |
|
9)
Message boards :
Number crunching :
Minirosetta v1.32 bug thread
(Message 54982)
Posted 7 Aug 2008 by netwraith
Post: Task reported too late to validate ????? Here are two tasks submitted by a new machine I just brought up. It's a slightly older 2.34 GLIBC machine using 5.8.16 for a client (All the newer ones want GLIBC 2.4+) The task was obtained, crunched and submitted the next day without any abort or error indicated in the STDOUT.. But with 0.00 credit ??? (The machine is only crunching Rosetta at this point). What is up with this ?? ... Could it be the machine ?? This is the first time I have seen this one... (and ... yes, the clocks are synchronized by NTP, so TOD should not be an issue) http://boinc.bakerlab.org/rosetta/result.php?resultid=182943687 stderr out <core_client_version>5.8.16</core_client_version> <![CDATA[ <stderr_txt> # cpu_run_time_pref: 43200 ====================================================== DONE :: 1 starting structures 43113 cpu seconds This process generated 84 decoys from 84 attempts ====================================================== BOINC :: Watchdog shutting down... BOINC :: BOINC support services shutting down... called boinc_finish </stderr_txt> ]]> Validate state Task was reported too late to validate Claimed credit 155.951746654641 Granted credit 0 application version 1.32 http://boinc.bakerlab.org/rosetta/result.php?resultid=182943208 stderr out <core_client_version>5.8.16</core_client_version> <![CDATA[ <stderr_txt> # cpu_run_time_pref: 43200 ====================================================== DONE :: 1 starting structures 42875.1 cpu seconds This process generated 123 decoys from 123 attempts ====================================================== BOINC :: Watchdog shutting down... BOINC :: BOINC support services shutting down... called boinc_finish </stderr_txt> ]]> Validate state Task was reported too late to validate Claimed credit 155.091236560251 Granted credit 0 application version 1.32 |
|
10)
Message boards :
Number crunching :
validator is down
(Message 54258)
Posted 7 Jul 2008 by netwraith
Post: Still says down... lots of stuff piling up... Maybe everyone is still out on the holiday weekend ??? |
|
11)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53811)
Posted 19 Jun 2008 by netwraith
Post: -- Mine are all Linux , but, several some above have been reported with Windows.. Most of my machines use 5.10.21 client , with 3 or 4 at 5.10.45.... The symptoms are the same for each client. I have already posted a half dozen work units to check and there have been a couple of dozen posted by others earlier in the thread... I would do the debug work for you, but, I don't have all the required code.... so... please read what others have posted... |
|
12)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53805)
Posted 18 Jun 2008 by netwraith
Post: -- The processes stay running, at or near 100% finished and continue until an abort, restart or deadline. Further, a restart will either finish without recrunching the last decoy or will start the last one over with (most likely) the same hang at or near 100%.. This is accomplished with the CPU/CORE index dropping to zero, but, with the condition that BOINC is unable to assign any other crunching job to the CPU/CORE. That core becomes useless, with the exception of handling system overhead. The watchdog process is either unable to correct the issue or fails to recognize that there is an issue. I will open up a system to more jobs (I had discontinued all my Rosetta crunching) to see if I can capture any more information about this. I agree that this would be a difficult one to track since the problem only shows in aborted jobs or, in the case of a successful restart intervention, a completed job. |
|
13)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53781)
Posted 18 Jun 2008 by netwraith
Post: My success ratio for restarts is now dropped to below 50%... I will now abort the hung units and will preemptively abort all t404 & t405 CASP8 units.... My apologies to Rosetta, but, I can't have my crunchers useless because of a bug in one project. Yes, I primarily crunch Rosetta, but..... |
|
14)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53744)
Posted 17 Jun 2008 by netwraith
Post: On the t404 & t405 CASP8 units that get stuck, there is a way to save the work. It's a bit of a pain, but, if you shutdown the connected client and then restart it, the task will finish on restart about 9 out of 10 times. The ones that don't finish will either continue from a percentage less than 100% or will self abort with a client error. Luck of the draw on this... I see no indicator as to why some units fail in this manner, but, it does save a majority of the units. This leads me to believe that there may be some sort of a watchdog issue on this. Why it would affect just those units is odd, but, this workaround may save a few hassles until the problem is found... |
|
15)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53724)
Posted 16 Jun 2008 by netwraith
Post: Looked at my multi-core AMD machine (64-bit Linux client) and saw that four cores were idle. In each case, the WU had "completed" (that is, had used up as much as it should of the time allocated), but had FAILED to tell the BOINC client that the task was done. As a result, the BOINC client kept that task dispatched on a CPU core -- but since the task had finished; the corresponding core was idle. [Boincmgr showed 100% complete for these tasks, and showed that these "running" tasks were NOT accumulating any more CPU time.] Unfortunately, the only way I had to get my CPUs running again was to abort these tasks (and that in turn caused their results to be thrown away -- my system had crunched these tasks uselessly). Here are a few of the aborts... All the machines with trouble are NetBurst or Dual Core systems... none of my uni-processor machines are showing these symptoms. The NetBurst MP systems have 4 or 8GB ram. The Dual-Cores are all 2GB ram.. so memory is probably not an issue. These all seem to be t404 or t405 CASP8 units.... http://boinc.bakerlab.org/rosetta/result.php?resultid=171381753 http://boinc.bakerlab.org/rosetta/result.php?resultid=171425636 http://boinc.bakerlab.org/rosetta/result.php?resultid=171446390 http://boinc.bakerlab.org/rosetta/result.php?resultid=171523734 http://boinc.bakerlab.org/rosetta/result.php?resultid=171536681 http://boinc.bakerlab.org/rosetta/result.php?resultid=171139128 |
|
16)
Message boards :
Number crunching :
Problems with version 5.96
(Message 53718)
Posted 16 Jun 2008 by netwraith
Post: Looked at my multi-core AMD machine (64-bit Linux client) and saw that four cores were idle. In each case, the WU had "completed" (that is, had used up as much as it should of the time allocated), but had FAILED to tell the BOINC client that the task was done. As a result, the BOINC client kept that task dispatched on a CPU core -- but since the task had finished; the corresponding core was idle. [Boincmgr showed 100% complete for these tasks, and showed that these "running" tasks were NOT accumulating any more CPU time.] Unfortunately, the only way I had to get my CPUs running again was to abort these tasks (and that in turn caused their results to be thrown away -- my system had crunched these tasks uselessly). Myself as well. I am seeing a lot of these. It does not seem to matter what the processor is, but, seems to be happening on the CASP8 tasks. Any reason that CASP8 is being run on a BETA version??? I will be aborting about a dozen units here in a few minutes. |
|
17)
Message boards :
Number crunching :
Minirosetta v1.28 bug thread
(Message 53694)
Posted 14 Jun 2008 by netwraith
Post: Some weirdness here.... Ran for 12 hours ... only one decoy generated... large disparity between claim and grant... What could be up with this ??? http://boinc.bakerlab.org/rosetta/result.php?resultid=170945330 Validate state Valid Claimed credit 114.084575208418 Granted credit 3.82141254947678 application version 1.28 |
|
18)
Message boards :
Number crunching :
Problems with version 5.90/5.91
(Message 50007)
Posted 24 Dec 2007 by netwraith
Post: Don't worry about the *CAUTION* I fat fingered an edit... Mea Culpa... -- |
|
19)
Message boards :
Number crunching :
Problems with version 5.90/5.91
(Message 50003)
Posted 24 Dec 2007 by netwraith
Post: -- *CAUTION* Since I originally posted this I had a few weird aborts. Maybe this should be ignored ... If I figure out exactly what happened I will post the issue *CAUTION* Another thing... while your client is down.... If you are adept with a text editor, you can edit the client_state.xml and change the 5.90 version references to 5.91. Just search for '590' It will find lines like: <version_num>590</version_num> for these change the 590 to 591. Then search for 5.90 It will find lines like: <file_name>rosetta_beta_5.90_i686-pc-linux-gnu</file_name> change the .90 to .91 One *CAVEAT* ::: the 5.90 search will also find lines like: <url>http://boinc.bakerlab.org/rosetta/download/rosetta_beta_5.90_i686-pc-linux-gnu</url> *Just leave these alone* Restart the client and 5.91 will be substituted for the 5.90... I don't know what would happen should you don this in the middle of a run, but, since, the alternative could be aborting the process..... I am not sure I would do it. What I would do if I were at Rosetta is re-run all of the Linux 5.90 processes just to be sure that the results are valid. I know about CRC's and checksums, but, I would rather not find out like Intel found out about the bug in the IEEE math processors of yore.... IIIIEEEEEEE!!!! Of course your mileage may vary... |
|
20)
Message boards :
Number crunching :
Problems with version 5.90/5.91
(Message 50002)
Posted 24 Dec 2007 by netwraith
Post: -- You don't need to abort the 5.90 runs... Just shutdown the client and restart. It will save the results.. Just do it after the normal time limit has expired and it won't restart the tasks... |
©2026 University of Washington
https://www.bakerlab.org