False -9 errors on older 9000 series GPU's.

Message boards : Number crunching : False -9 errors on older 9000 series GPU's.
Message board moderation

To post messages, you must log in.

AuthorMessage
Profile Fred J. Verster
Volunteer tester
Avatar

Send message
Joined: 21 Apr 04
Posts: 3252
Credit: 31,903,643
RAC: 0
Netherlands
Message 1084315 - Posted: 6 Mar 2011, 13:39:35 UTC
Last modified: 6 Mar 2011, 13:49:39 UTC

This false -9,(SETI@Home Informational message -9 result_overflow
NOTE: The number of results detected exceeds the storage space allocated.

Flopcounter: 334319333.801968

Spike count: 11
Pulse count: 20
Triplet count: 0
Gaussian count: 0
called boinc_finish.

I've seen alot of them on my 9800GTX+, in the past, heat was in my case the cause, replaced it with an GTS250 (in fact the 'same', G92), but less succeptable to heat, cause it didn't make those faults.
(Wonder how much of these errors, were made in the past and slipped through,
when a canonnical result was made with 2 -9 results?)

Although most canonnical results, came from a CPU vs a GPU.
But we have better Software on better and faster Hardware, nowadays!
ID: 1084315 · Report as offensive
Richard Haselgrove Project Donor
Volunteer tester

Send message
Joined: 4 Jul 99
Posts: 14690
Credit: 200,643,578
RAC: 874
United Kingdom
Message 1084319 - Posted: 6 Mar 2011, 13:53:13 UTC - in response to Message 1084315.  

The false -9 on older cards has been observed in the past. It can be overheating, memory corruption following an application crash, and probably a number of other types of problem as well. I've seen it a couple of times on my 9800-series cards. So far as I know, it can only be cleared by restarting the host computer, but they usually come back to life and crunch normally once that has been done.

Because these crashes are random, and depend on some combination of problem and circumstance unique to the card, I don't think there's much danger of two such cases validating each other.

That's different from the deliberate case, where Fermi-class GPUs are loaded with an incompatible app (Raistmer's V12 - don't do it!). Then, because both -9s are created the same way, they are all too likely to validate - only took me a few seconds to find WU 706601640 - two of our old friends from the list Joe and Claggy drew up months ago, still at it.
ID: 1084319 · Report as offensive
Profile Fred J. Verster
Volunteer tester
Avatar

Send message
Joined: 21 Apr 04
Posts: 3252
Credit: 31,903,643
RAC: 0
Netherlands
Message 1084325 - Posted: 6 Mar 2011, 14:20:20 UTC - in response to Message 1084319.  

A 'nice' also clear example of using the wrong app.!

Poor wingman:Device 1 : GeForce GT 240
totalGlobalMem = 1034092544
sharedMemPerBlock = 16384
regsPerBlock = 16384
warpSize = 32
memPitch = 2147483647
maxThreadsPerBlock = 512
clockRate = 1340000
totalConstMem = 65536
major = 1
minor = 2
textureAlignment = 256
deviceOverlap = 1
multiProcessorCount = 12
setiathome_CUDA: CUDA Device 1 specified, checking...
Device 1: GeForce GT 240 is okay
SETI@home using CUDA accelerated device GeForce GT 240
setiathome_enhanced 6.09 Visual Studio/Microsoft C++
libboinc: 6.3.22

Work Unit Info:
...............
WU true angle range is : 0.410107
Optimal function choices:
-----------------------------------------------------
name
-----------------------------------------------------
v_BaseLineSmooth (no other)
v_GetPowerSpectrum 0.00035 0.00000
v_ChirpData 0.02050 0.00000
v_Transpose4 0.01425 0.00000
FPU opt folding 0.00400 0.00000

Flopcounter: 42152786330206.180000

Spike count: 6
Pulse count: 4
Triplet count: 0
Gaussian count: 2
called boinc_finish

who likely has delivered the valid result and gets 'ruled out' ?!
I suppose these are being Resend..



ID: 1084325 · Report as offensive
Profile Link
Avatar

Send message
Joined: 18 Sep 03
Posts: 834
Credit: 1,807,369
RAC: 0
Germany
Message 1084328 - Posted: 6 Mar 2011, 14:23:07 UTC - in response to Message 1084319.  

Then, because both -9s are created the same way, they are all too likely to validate - only took me a few seconds to find WU 706601640 - two of our old friends from the list Joe and Claggy drew up months ago, still at it.

Should the new quota system not make it more difficult for such hosts to get WUs for the faulty device? However both of them have a quota of 100 for the GPUs, so that's still not working.
ID: 1084328 · Report as offensive
Profile perryjay
Volunteer tester
Avatar

Send message
Joined: 20 Aug 02
Posts: 3377
Credit: 20,676,751
RAC: 0
United States
Message 1084365 - Posted: 6 Mar 2011, 16:35:14 UTC

I've also been seeing a few 9xxx series cards giving -9s. I usually don't bother with them as I figure they will have to reboot sometime and that will probably cure most of them. If, on the other hand it is heat or bad card related that is something I couldn't tell them in a short private message. Another thing I'm seeing is half of a 295 card throwing -9s. I guess that is probably heat related or bad card too.

I'm also seeing some familiar names with the V12 App on a Fermi card. I have tried PMing them but it seems to do little good. Either they have notification turned off or they just don't care. I think I've only had one person reply thanking me and correcting the problem.


PROUD MEMBER OF Team Starfire World BOINC
ID: 1084365 · Report as offensive
Claggy
Volunteer tester

Send message
Joined: 5 Jul 99
Posts: 4654
Credit: 47,537,079
RAC: 4
United Kingdom
Message 1084369 - Posted: 6 Mar 2011, 16:43:04 UTC - in response to Message 1084319.  

Perhaps we should contact Eric again, perhaps as a Project Admin he can turn off GPU requests for those hosts/users,

Claggy
ID: 1084369 · Report as offensive
Profile HAL9000
Volunteer tester
Avatar

Send message
Joined: 11 Sep 99
Posts: 6534
Credit: 196,805,888
RAC: 57
United States
Message 1084406 - Posted: 6 Mar 2011, 18:20:27 UTC

Whenever I have a task "Completed, validation inconclusive" it is always against a GPU that sent it back in with -9. Looking them over the majority of them are 9000 cards with 400 cards coming up 2nd. The 200 cards seem to only spit out -9 results when a CPU does also. Albeit slower than a CPU.

GPU processing hasn't had 40+ years of development behind it, but I think 3DFX, hey remember them, had a point.


SETI@home classic workunits: 93,865 CPU time: 863,447 hours
Join the [url=http://tinyurl.com/8y46zvu]BP6/VP6 User Group[
ID: 1084406 · Report as offensive
Profile Westsail and *Pyxey*
Volunteer tester
Avatar

Send message
Joined: 26 Jul 99
Posts: 338
Credit: 20,544,999
RAC: 0
United States
Message 1084501 - Posted: 6 Mar 2011, 23:59:27 UTC

Here is one to keep an eye on:

Just noticed this WU...Mine is the first CPU result..
wuid=695733638

???
"The most exciting phrase to hear in science, the one that heralds new discoveries, is not Eureka! (I found it!) but rather, 'hmm... that's funny...'" -- Isaac Asimov
ID: 1084501 · Report as offensive

Message boards : Number crunching : False -9 errors on older 9000 series GPU's.


 
©2026 University of California
 
SETI@home and Astropulse are funded by grants from the National Science Foundation, NASA, and donations from SETI@home volunteers. AstroPulse is funded in part by the NSF through grant AST-0307956.