Dark Bit Factory & Gravity
PROGRAMMING => Other languages => Blitz => Topic started by: Devils Child on July 09, 2007
-
did you ever seen the hl1 software renderer?
on a screen resoultion of 1280x1024 i get the full 75 FPS!
and when i make a simple for..next loop in bb where i fill the whole screen i just get 40 FPS
something's wrong there...
half-life gets full FPS and i only get 40 fps on a much simmpler fill screen routine.
what is hl doing different than i?
-
"what is hl doing different than i?"
You'd probably have to ask Abrash or Carmack.
I always had problems getting blitz to do much in anything over 320*240 other than maybe a couple of objects or something but certainly not the larger scenes with some effects going on that I managed at 320*240, that was a couple of years ago and on a lower spec machine than I have now but moving over to freebasic at the time made 640*480 a bit more of a realistic possibility.
Now on to using c++, I think 640*480 is fine speedwise but 800*600 is still pushing it a bit, not to say it can't be done other ways but it's still a bit beyond me.
HL was very heavily optimised from the type of maps they used right down to the low level code so it comes down to a lot of different things.
-
hmh, since i found out that blitz3d is totaly dis-optimized, i'm thinking of converting my virtualGL lib to bmax and test it there.
i think i'll learn it within 3 days. (dont worry, i learned php within a week)
hm, i'm so jelous of carmacs crew. they made a software renderer of top quality. also they made doom3. it is using multiple stencil lights and lots of shaders, and it runs on 75 FPS at all the times without getting under 60 fps
so jelous ...
-
Mmm, I have to echo your sentiment there and also Stonemonkeys.. I used to use Blitz a lot but got totally disenchanted with it because of the massive file sizes and in particular the slow render speed of writepixelfast. I did write one or two demos with texture mapping in 640 X 480, but I had to use all sorts of tricks to get them up to speed.
In my experience anything over 640 X 480 was a waste of time in Blitz. You will certainly be better served with another language and with someone of your obvious ability I am sure you'll have a ball when you have some proper power to play with.
Btw Stonemonkey, I might be far off the beam here but I think I read somewhere that there are some compiler options you can use to increase the execution speed of VC++ and I'm sure it was Jim who posted it.
-
yeah you will be talking about turning the sse1/2 switches on in visual studio that can make a huge diffrence in speed also when you switch to release mode that makes a big speed difrence as well.
-
Hmm, I tried the SSE options but it ended up being slower :( .
-
really, i havent yoused it much in visual studio but in dev c i get between 7 and 10x speed increases doing perspective correct mapping and matrix ops setting in up to generate code for a pentium 4 sse2.
-
my hope were some pixel dlls like this:#
http://dbfinteractive.com/index.php?topic=1121.0
but this is 4 times slower than blitz!
the reasong might be, that for each pixel i have to call a function in a dll...
writepixelfast writes some bytes to a buffer of an image, the screen, a texture or whatever.
is it faster to write it directly using PokeInt? and if, how?
-
There is a way to access the screen memory as a bank in B+, but not sure if that's the case for BB or B3D but what I usually did was to render everything into a 1D array of ints then copy that to the screenbuffer using writepixelfast once the scene was drawn.
-
This is the best i can find on the BB site:
http://www.blitzbasic.com/codearcs/codearcs.php?code=1104
//Paul
-
hm this sounds cool. i tried it.
[...]
thank you
-
I managed to get about 25fps out of blitz at 800x600 with a wbuffer filling about 2/3 screen, but it was so stupid. Everything was hampered by the 'copy to screen and clear the wbuffer' stuff, which took nearly all the time.
If you were writing halflife, you wouldn't bother trying to do any of that stuff at high level, you'd dive straight in at the 'write direct to memory' level. You should be able to hammer several hundred frames a second over to a modern video card, so do all the work in system RAM and then blast the finished frames over using an optimised copy.
Jim
-
hmh. i've seen the pixel-to-RAM code above, but it is slower than writepixelfast on my machine and faster than writepixelfast on an other machine.
also on my machine it is both at same speed using 320x240 and when you use 1280x1024 writepixelfast is 5x faster?!?
so what should i take?
-
Totally forget writepixelfast. It's useless because each time you write a pixel you supply the x and y coordinate, so to write a pixel it has to calculate:
address = screen_base_address + (x + y * screen_width) * sizeof_pixel
This is futile, since more than likely the last pixel you drew was at x-1 (the previous pixel just next to the current one) and you could get its address by just adding sizeof_pixel (usually 4 for a 32bit screen) to the previous address! This is how you can get far more speed by going in at a lower level.
Best thing to do is to write to a system memory buffer and use memcpy() to move the data up to the screen in one go, that will be within 10% of optimum.
By the way 1280x1024x32@75fps is a smidge under 400Megabytes/second. A modern PCI Express card can in theory do 16x that speed. Reading from the screen is far, far slower.
Jim
-
this is a test i made:
.lib "kernel32.dll"
apiRtlMoveMemory(Destination*,Source,Length):"RtlMoveMemory"
apiRtlMoveMemory2(Destination,Source*,Length):"RtlMoveMemory"
Const gw = 1024
Const gh = 768
Const gw1 = gw - 1
Const gh1 = gh - 1
Const gw2 = gw / 2 - 1
Const gw21 = gw / 2
Graphics gw, gh, 32, 2
SetBuffer BackBuffer()
bnkVideo = CreateBank(gw * gh * 4)
w4 = gw * 4
start = MilliSecs()
For i = 1 To 30
For x = 0 To gw2
For y = 0 To gh1
PokeInt bnkVideo, y * w4 + x * 4, $00FF00
Next
Next
Next
LockBuffer()
bnkInfo = CreateBank(32)
apiRtlMoveMemory bnkInfo, GraphicsBuffer() + 72, 32
size = PeekInt(bnkInfo, 20) * PeekInt(bnkInfo, 24) * PeekInt(bnkInfo, 28) / 8
If BankSize(bnkVideo) < size Or PeekInt(bnkInfo, 0) = 0 Then
FreeBank bnkInfo
RuntimeError "Failed to draw on buffer."
Else
apiRtlMoveMemory2 PeekInt(bnkInfo, 0), bnkVideo, size
FreeBank bnkInfo
EndIf
UnlockBuffer()
time1 = MilliSecs() - start
LockBuffer()
start = MilliSecs()
For i = 1 To 30
For x = gw21 To gw1
For y = 0 To gh1
WritePixelFast x, y, $FF0000
Next
Next
Next
time2 = MilliSecs() - start
UnlockBuffer()
Text 10, 10, "Poke: " + time1
Text gw2 + 10, 10, "WritePixelFast: " + time2
Flip
WaitKey()
End
on my machine, writepixelfast is 4 times faster on high resolutions...
why ain't i just writing >>DIRECTLY<< to the backbuffer? i mean - this code is modifyed of the link above
so on some machines the poke method is faster and on some machines the pixel method is faster. also on high resolutions the writepixel method is faster...
-
You're not going to get the best speed with blitz. Both ways are going to be too slow.
Some suggestions though.
Try changing the PokeInt loop a bit
For x = 0 To gw2
For y = 0 To gh1
PokeInt bnkVideo, y * w4 + x * 4, $00FF00
Next
Next
to
size%=gh1*gw2-1
address%= 0
for c%= 0 to size
pokeint bnkVideo,address,$00ff00
address = address+4
next
It's still not going to be fast, but it omits a heap of extra multiplication PER PIXEL which you must avoid at all costs.
Also, try moving the CreateBank/FreeBank outside the loop. I suspect the allocation is pretty slow.
The (obvious) problem though is here you're comparing apples with oranges. The bank version is writing to system memory, reading back from system memory, then writing to video memory - the wpf version is writing direct to video memory. Clearly even with the hideous overheads of wpf it's going to be faster. The point is supposed to be that writing to system memory is supposed to be so quick that the extra write/read is hidden. Problem is PokeInt is too slow. Also, this isn't a real life scenario. Often you write each pixel more than once, and read it back to do alpha effects. If you tried that with ReadPixelFast it would be a disaster, with PeekInt it's much more possible.
Given the theoretical maximum you can blast across to the video card is 4GB/s, can you work out roughly how much bandwidth you're getting from each of these 2 methods to see how far off theoretical it is? That would be pretty interesting.
<edit>Just realised you're writing your pixels in vertical columns. That's really bad for the CPU cache. Much better to write in horizontal rows where possible.
Jim
-
hm, semms like i better trust you in that ;)
why is your software renderer using wpf then?
i cant implement vertical colums anymore. it would be to difficult to manage...
thank you for these fast replies :)
-
here is my (slower) vgl lib using the poke method
http://dc.freecoder-portal.de/upload/files/Devils%20Child/VirtualGL.zip
am i doing everything right?
-
i cant implement vertical colums anymore. it would be to difficult to manage...
Jim was meaning that in your tests you have the y loop as the inner loop which fills the buffer column by column which is bad for the cache.
-
you mean this one?
mi = GraphicsWidth() * GraphicsHeight() * 4
i = 0
While i < mi
PokeInt vglScreenBank, i, col
i = i + 4
Wendthis is as optimized as it could be...
-
In the loops like this:
For x = 0 To gw2
For y = 0 To gh1
PokeInt bnkVideo, y * w4 + x * 4, $00FF00
Next
Next
where the y loop is the inner loop.
-
this is faster anyway:
mi = GraphicsWidth() * GraphicsHeight() * 4
i = 0
While i < mi
PokeInt vglScreenBank, i, col
i = i + 4
Wend
i think jim really meant my fill triangle routine, where i should take vertical lines instead of horizontals...
but whatever, i cant do that because it would mean to completely redo the triangle fill routine
is it really faster using vertical lines instead of my horizontal? and if, how much faster is it?
-
I think Jim meant what Stonemonkey said.
Anyway, lots of people have said it here, I'll say it again.
You have outgrown Blitz. It is not capable of the sort of speedy pixel manipulation you need, no matter how you try and do it you need to learn another language.
I wrote two quick comparisons for you, both run in 800 * 600 mode (see attached zip for exes).
FPS is displayed in the console window (you may have to move the display window out of the way).
Firstly, if you were rasterising a solid colour you could use something like this;
' FAST PIXELWISE BLITTING USING SIMPLE RASTERISATION.
'
'---------------------------------------------------------------------------
#DEFINE PTC_WIN
#INCLUDE "TINYPTC.BI"
#INCLUDE "WINDOWS.BI"
OPTION STATIC
OPTION EXPLICIT
CONST XRES = 800
CONST YRES = 600
DIM SHARED AS UINTEGER BUFFER ( XRES * YRES )
DIM SHARED AS INTEGER Y,SLICE,TC,TICKS
DIM SHARED AS UINTEGER PTR PP
DIM SHARED AS DOUBLE GRAB
IF( PTC_OPEN( "SPEED TEST", XRES, YRES ) = 0 ) THEN
END -1
END IF
GRAB=TIMER
WHILE(GETASYNCKEYSTATE(VK_ESCAPE)<>-32767)
FOR Y=0 TO YRES-1
SLICE = XRES
TC = RGB(Y /4 , Y/3 , Y/5 )
PP = @BUFFER(Y*XRES)
asm
mov eax,dword ptr[TC]
mov ecx, [slice]
mov edi, [PP]
rep stosd
end asm
NEXT
PTC_UPDATE@BUFFER(0)
TICKS=TICKS+1
IF TIMER-GRAB>=1 THEN
PRINT "FPS : " + STR(TICKS)
TICKS=0
GRAB=TIMER
END IF
WEND
And to pick pixels out of a work buffer, it can be done as fast as this;
' FAST PIXELWISE BLITTING USING POINTERS.
'
'---------------------------------------------------------------------------
#DEFINE PTC_WIN
#INCLUDE "TINYPTC.BI"
#INCLUDE "WINDOWS.BI"
OPTION STATIC
OPTION EXPLICIT
CONST XRES = 800
CONST YRES = 600
DIM SHARED AS UINTEGER BUFFER ( XRES * YRES )
DIM SHARED AS INTEGER Y,SLICE,TC,TICKS,X
DIM SHARED AS UINTEGER PTR PP
DIM SHARED AS DOUBLE GRAB
IF( PTC_OPEN( "SPEED TEST", XRES, YRES ) = 0 ) THEN
END -1
END IF
GRAB=TIMER
WHILE(GETASYNCKEYSTATE(VK_ESCAPE)<>-32767)
PP=@BUFFER(0)
FOR Y=0 TO YRES * XRES -1
*PP = Y
PP+=1
NEXT
PTC_UPDATE@BUFFER(0)
TICKS=TICKS+1
IF TIMER-GRAB>=1 THEN
PRINT "FPS : " + STR(TICKS)
TICKS=0
GRAB=TIMER
END IF
WEND
Or just invoke memcpy if you want to include crt.bi.
In any case, this is in another variant of Basic, you would find if fairly easy to port your code over and as you'll see from the axamples it runs about 6-8 times faster than Blitz on most computers for pixelwise stuff.
If you are software rendering, you can persist with Blitz but in the end, trust me, you will eventually get pissed off and switch to another language.
-
Just realised you're writing your pixels in vertical columns. That's really bad for the CPU cache. Much better to write in horizontal rows where possible.
It is much better to write with horizontal lines which makes better use of the cache (in any language). As I mentioned before, the tests you were doing to fill the screenbuffer had y for the inner loop which meant it was drawing in columns which is not as good.
-
why is your software renderer using wpf then?
I use banks/pokeint to do all the drawing, then wpf only the pixels that were written (by checking the wbuffer) to the screen. Can you believe it's actually faster to do an if test for every pixel than it is to draw the pixel?! And because I'd never seen this rtl stuff :)
The code's here if you want to play with it.
http://members.iinet.net.au/~jimshaw/Blitz/typetriw.zip (http://members.iinet.net.au/~jimshaw/Blitz/typetriw.zip)
Jim
-
(http://www.dc.freecoder-portal.de/upload/files/Devils%20Child/Screen169.jpg)
hmh. most of you would recognize this as the first DOOM level. except for the second screenshot...
this is written using the poke method.
1. all resolutions work with windowed mode
2. all resolutions work in fullscreen EXCEPT 400x300x32x2
-
That 2nd shot looks like it's a problem with the pitch, that the memory set aside for the screen is more than is required for the actual screen. Even though the visible screen is 400*300 the buffer might actually be 640*300 (or something else) with the left hand side not visible.
-
I'm surprised the driver/card actually accepts 400x300 as a fullscreen resolution. It's way off standard.
Jim