Message boards :
Number crunching :
False -9 errors on older 9000 series GPU's.
Message board moderation
| Author | Message |
|---|---|
Fred J. Verster Send message Joined: 21 Apr 04 Posts: 3252 Credit: 31,903,643 RAC: 0
|
This false -9,(SETI@Home Informational message -9 result_overflow NOTE: The number of results detected exceeds the storage space allocated. Flopcounter: 334319333.801968 Spike count: 11 Pulse count: 20 Triplet count: 0 Gaussian count: 0 called boinc_finish. I've seen alot of them on my 9800GTX+, in the past, heat was in my case the cause, replaced it with an GTS250 (in fact the 'same', G92), but less succeptable to heat, cause it didn't make those faults. (Wonder how much of these errors, were made in the past and slipped through, when a canonnical result was made with 2 -9 results?) Although most canonnical results, came from a CPU vs a GPU. But we have better Software on better and faster Hardware, nowadays!
|
Richard Haselgrove ![]() Send message Joined: 4 Jul 99 Posts: 14690 Credit: 200,643,578 RAC: 874
|
The false -9 on older cards has been observed in the past. It can be overheating, memory corruption following an application crash, and probably a number of other types of problem as well. I've seen it a couple of times on my 9800-series cards. So far as I know, it can only be cleared by restarting the host computer, but they usually come back to life and crunch normally once that has been done. Because these crashes are random, and depend on some combination of problem and circumstance unique to the card, I don't think there's much danger of two such cases validating each other. That's different from the deliberate case, where Fermi-class GPUs are loaded with an incompatible app (Raistmer's V12 - don't do it!). Then, because both -9s are created the same way, they are all too likely to validate - only took me a few seconds to find WU 706601640 - two of our old friends from the list Joe and Claggy drew up months ago, still at it. |
Fred J. Verster Send message Joined: 21 Apr 04 Posts: 3252 Credit: 31,903,643 RAC: 0
|
A 'nice' also clear example of using the wrong app.! Poor wingman:Device 1 : GeForce GT 240 totalGlobalMem = 1034092544 sharedMemPerBlock = 16384 regsPerBlock = 16384 warpSize = 32 memPitch = 2147483647 maxThreadsPerBlock = 512 clockRate = 1340000 totalConstMem = 65536 major = 1 minor = 2 textureAlignment = 256 deviceOverlap = 1 multiProcessorCount = 12 setiathome_CUDA: CUDA Device 1 specified, checking... Device 1: GeForce GT 240 is okay SETI@home using CUDA accelerated device GeForce GT 240 setiathome_enhanced 6.09 Visual Studio/Microsoft C++ libboinc: 6.3.22 Work Unit Info: ............... WU true angle range is : 0.410107 Optimal function choices: ----------------------------------------------------- name ----------------------------------------------------- v_BaseLineSmooth (no other) v_GetPowerSpectrum 0.00035 0.00000 v_ChirpData 0.02050 0.00000 v_Transpose4 0.01425 0.00000 FPU opt folding 0.00400 0.00000 Flopcounter: 42152786330206.180000 Spike count: 6 Pulse count: 4 Triplet count: 0 Gaussian count: 2 called boinc_finish who likely has delivered the valid result and gets 'ruled out' ?! I suppose these are being Resend..
|
Link Send message Joined: 18 Sep 03 Posts: 834 Credit: 1,807,369 RAC: 0
|
Then, because both -9s are created the same way, they are all too likely to validate - only took me a few seconds to find WU 706601640 - two of our old friends from the list Joe and Claggy drew up months ago, still at it. Should the new quota system not make it more difficult for such hosts to get WUs for the faulty device? However both of them have a quota of 100 for the GPUs, so that's still not working.
|
perryjay Send message Joined: 20 Aug 02 Posts: 3377 Credit: 20,676,751 RAC: 0
|
I've also been seeing a few 9xxx series cards giving -9s. I usually don't bother with them as I figure they will have to reboot sometime and that will probably cure most of them. If, on the other hand it is heat or bad card related that is something I couldn't tell them in a short private message. Another thing I'm seeing is half of a 295 card throwing -9s. I guess that is probably heat related or bad card too. I'm also seeing some familiar names with the V12 App on a Fermi card. I have tried PMing them but it seems to do little good. Either they have notification turned off or they just don't care. I think I've only had one person reply thanking me and correcting the problem. PROUD MEMBER OF Team Starfire World BOINC |
|
Claggy Send message Joined: 5 Jul 99 Posts: 4654 Credit: 47,537,079 RAC: 4
|
Perhaps we should contact Eric again, perhaps as a Project Admin he can turn off GPU requests for those hosts/users, Claggy |
HAL9000 Send message Joined: 11 Sep 99 Posts: 6534 Credit: 196,805,888 RAC: 57
|
Whenever I have a task "Completed, validation inconclusive" it is always against a GPU that sent it back in with -9. Looking them over the majority of them are 9000 cards with 400 cards coming up 2nd. The 200 cards seem to only spit out -9 results when a CPU does also. Albeit slower than a CPU. GPU processing hasn't had 40+ years of development behind it, but I think 3DFX, hey remember them, had a point. SETI@home classic workunits: 93,865 CPU time: 863,447 hours Join the [url=http://tinyurl.com/8y46zvu]BP6/VP6 User Group[
|
Westsail and *Pyxey* Send message Joined: 26 Jul 99 Posts: 338 Credit: 20,544,999 RAC: 0
|
Here is one to keep an eye on: Just noticed this WU...Mine is the first CPU result.. wuid=695733638 ??? "The most exciting phrase to hear in science, the one that heralds new discoveries, is not Eureka! (I found it!) but rather, 'hmm... that's funny...'" -- Isaac Asimov
|
©2026 University of California
SETI@home and Astropulse are funded by grants from the National Science Foundation, NASA, and donations from SETI@home volunteers. AstroPulse is funded in part by the NSF through grant AST-0307956.