|
1)
Message boards :
Number crunching :
High write-rate with Pythons
(Message 106866)
Posted 5 Sep 2022 by Paddles Post: When I included a write-cache (it was 16 GB with a 4 hour latency), the writes to the SSD itself dropped to a negligible amount of less than about 1 GB/day. So the problem is solved for me. To follow up... I did eventually install another SSD to serve solely for BOINC (might end up doubling as scratch space for anything else that needs it), on the basis that I'd rather lose a secondary SSD than the one that holds Windows and all my software. I've also installed PrimoCache with caching just for that drive (16GB out of 32GB, which should be enough to handle my configured maximum of 4 VBOX jobs, and defer-write set to Infinite). So far the caching seems to be reducing the amount of data actually written to the SSD to about 40% of what it would otherwise be, and the amount of data written to my primary drive has reduced significantly. It looks like some BOINC jobs complete and and get cleaned up using data in cache, resulting in little or no data being written to disk. If/when the BOINC/scratch disk dies, I'll just replace it with another low-end SSD, but with PrimoCache as part of the mix this could be 2-3 years away - and who knows what R@H will be doing then? So thanks to jim and others for suggestions/guidance. |
|
2)
Message boards :
Number crunching :
High write-rate with Pythons
(Message 106710)
Posted 4 Aug 2022 by Paddles Post: When I included a write-cache (it was 16 GB with a 4 hour latency), the writes to the SSD itself dropped to a negligible amount of less than about 1 GB/day. So the problem is solved for me. I am running a Windows machine - and given that BOINC data is on the main Windows system disk/volume (OS and applications; most of my non-BOINC data is on a separate SSD) I'd rather lose a sacrificial disk that's only used for BOINC than the disk that keeps everything going. Adding caching is an interesting idea, but not sure how well it would work on the main system disk. |
|
3)
Message boards :
Number crunching :
High write-rate with Pythons
(Message 106701)
Posted 3 Aug 2022 by Paddles Post: This discussion, and noticing the high amount of data written to my drive, had me wondering: would getting a separate SSD to use solely for BOINC data be sensible, for risk-management (to reduce the wear rate on the primary system SSD)? Or is that overly paranoid? My computer is not quite 1 year old, the 500GB SSD has 131TB written. I don't know how much of that is R@H VBox jobs but I'm guessing a lot of it. |
|
4)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 106118)
Posted 29 Apr 2022 by Paddles Post: I may have spoken too soon. The tasks were running for exceptionally long times (18-26 hours) - although unlike the normal "not doing anything" vbox tasks, they were showing significant CPU time utilised (rather than the tasks that "run" for 18 hours but have only consumed 10-20 seconds of CPU). I shut down BOINC, rolled VBox back to version 6.1.12 (BOINC recommended version, not 6.1.32 which I had been running), restarted, and all the vbox tasks came up with computation errors.I'm on VB version 5 (or Cosmology breaks completely) and it seems to run Rosetta Python just as well as 6. LHC is also happy with it. Kryptos at Home also hates 6. Seems like VB screwed up when they made the new one. Downgraded to VBox 5, everything was happy. Tried upgrading back to VBox 6.1.12, and everything continued to be happy, including an in-progress task resuming and completing normally. So maybe there was something up with 6.1.34 - and in-progress tasks can cope with VBox upgrades but not downgrades. So looks like I'll stick with VBox 6.1.12 for the time being. I'm only using it for R@H (or other enrolled BOINC projects if they can make use of it) so not being up-to-date won't affect anything else. |
|
5)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 106028)
Posted 25 Apr 2022 by Paddles Post: Update: The first task to be postponed reached the end of its one day postponement, and now appears to be computing successfully (in VBox 6.1.34). Haven't tried reverting to previous version to see what happens, but whatever the problem was it seems to have resolved. I may have spoken too soon. The tasks were running for exceptionally long times (18-26 hours) - although unlike the normal "not doing anything" vbox tasks, they were showing significant CPU time utilised (rather than the tasks that "run" for 18 hours but have only consumed 10-20 seconds of CPU). I shut down BOINC, rolled VBox back to version 6.1.12 (BOINC recommended version, not 6.1.32 which I had been running), restarted, and all the vbox tasks came up with computation errors. Oh well, will see what happens with the next tasks to run. |
|
6)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 106007)
Posted 24 Apr 2022 by Paddles Post: Update: The first task to be postponed reached the end of its one day postponement, and now appears to be computing successfully (in VBox 6.1.34). Haven't tried reverting to previous version to see what happens, but whatever the problem was it seems to have resolved. |
|
7)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 106006)
Posted 24 Apr 2022 by Paddles Post: Can you unupdate to 5.2.44? I'll give that a try if there are continued problems. I just went back and looked, the existing tasks are still postponed (1 day hasn't elapsed yet, and there isn't an obvious way to manually resume them), but one of the other python tasks seems to be running ok so maybe it was just a transient issue. |
|
8)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 106003)
Posted 24 Apr 2022 by Paddles Post: I'm encountering a new problem, has anyone else see it?. I have three Python tasks in the state "Postponed: VM Hypervisor failed to enter an online state in a timely fashion." I'm running BOINC 7.6.20 and was on VirtualBox 6.1.32 - combination had been working generally happy, and I haven't changed any BOINC settings recently (or had any significant changes to available disk space). I've updated VBox to 6.1.34 to see if that resolves it, but it looks like the Python tasks are being postponed for a day and none of the others have started. |
|
9)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 103878)
Posted 22 Dec 2021 by Paddles Post: I too was having problems with too many of the vbox jobs trying to run at once, leaving not enough resources resulting in Rosetta and other project tasks not being able to run, resulting in lower overall performance. In my case, lowering Rosetta's resource from 200 (out of 500 - 40% share) to 100 (out of 400 - 25% share) is giving me 10-15% more work completed across all projects. I haven't had any vbox jobs that had to be aborted since then. BOINC is allowed 11/12 cores, 75% of 32GB RAM. Other projects are WCG and Yoyo. |
|
10)
Message boards :
News :
Thank you!
(Message 103864)
Posted 19 Dec 2021 by Paddles Post: I was having problems with too many of the vbox jobs trying to run at once, leaving not enough resources resulting in Rosetta and other project tasks not being able to run, resulting in lower overall performance. In my case, lowering Rosetta's resource from 200 (out of 500 - 40% share) to 100 (out of 400 - 25% share) is giving me 10-15% more work completed across all projects, haven't had any vbox jobs that had to be aborted since then. BOINC is allowed 11/12 cores, 75% of 32GB RAM. Other projects are WCG and Yoyo. |
|
11)
Message boards :
Number crunching :
Problems and Technical Issues with Rosetta@home
(Message 103101)
Posted 4 Nov 2021 by Paddles Post: A task running MUCH longer than the expected 8 hours: Reassuring it's not just me. I've had a couple of the vbox tasks do that. I just aborted a task that had been running for 2d 12 hours elapsed, but only about 5 minutes of CPU. Supposedly was 99.8% complete but I think it was "99.6% complete" a day ago, and has gone a past deadline. (https://boinc.bakerlab.org/rosetta/result.php?resultid=1443593989 In this situation, is aborting it the most useful thing to do? I'm not really worried about credits or losing them - just want the CPU time to be doing something useful. Is letting it go on long past deadline still useful to someone, or should I manually abort tasks that seem to be wandering aimlessly so that the processing slot can go to another task? |
©2026 University of Washington
https://www.bakerlab.org