Hi Garry,
I've hit a performance wall and would like to share measurements and two requests.
1) Observation: every loop iteration costs about 25 µs
A query whose SQL runs in 25 ms on the server takes about 3.6 s end to end in Objo. The server response is only 1,560 bytes compressed / 7,691 bytes raw. Timing breakdown:
inflate (pure Objo DEFLATE decoder) 2,889 ms
Adler-32 over 7,691 bytes 289 ms
deserialize 7,691 bytes 532 ms
To find the cause I benchmarked individual operations:
Sub BenchPrimitives()
Var sizes() As Integer = [1000, 2000, 4000, 8000]
For Each n As Integer In sizes
Var mb As New MemoryBlock(n)
Var arr() As Integer
Var x As Integer = 0
Var t0 As Double = System.TicksMilliseconds
For i As Integer = 0 To n - 1
x = x + mb.ReadByte(i)
Next i
Var t1 As Double = System.TicksMilliseconds
For i As Integer = 0 To n - 1
mb.WriteByte(i, 7)
Next i
Var t2 As Double = System.TicksMilliseconds
For i As Integer = 0 To n - 1
arr.Append(i)
Next i
Var t3 As Double = System.TicksMilliseconds
For i As Integer = 0 To n - 1
x = x + arr[i]
Next i
Var t4 As Double = System.TicksMilliseconds
For i As Integer = 0 To n - 1
x = x + (i Mod 7)
Next i
Var t5 As Double = System.TicksMilliseconds
Print(n.ToString() + ": ReadByte " + (t1 - t0).ToString("F1") _
+ " | WriteByte " + (t2 - t1).ToString("F1") _
+ " | Append " + (t3 - t2).ToString("F1") _
+ " | arr[i] " + (t4 - t3).ToString("F1") _
+ " | Mod " + (t5 - t4).ToString("F1") + " ms")
Next n
End Sub
Results (ms):
n ReadByte WriteByte Append arr Mod
1000 77.4 40.1 38.6 36.6 40.6
2000 72.3 72.3 78.2 66.0 51.2
4000 101.0 101.0 100.8 101.8 102.1
8000 201.6 201.4 201.5 203.1 203.7
Everything scales linearly, but every column costs the same 25 µs per iteration regardless of the operation - even a plain x = x + (i Mod 7). That suggests the cost is per-iteration/per-statement overhead rather than in the operations themselves (roughly 40,000 loop iterations per second).
An empty loop For i As Integer = 0 To 7999 : Next i takes 27 ms (about 3.4 µs per iteration), so the loop itself is cheap. Adding a single statement to the body raises it to 200 ms, meaning each statement costs roughly 22 µs regardless of what it does. That looks like a fixed per-statement overhead, possibly debugger line tracking.
Is this expected, or could something like debugger instrumentation or event pumping be running on every loop iteration?
2) Request: native zlib / DEFLATE compression
B4XSerializator wraps every payload in a zlib stream (java.util.zip.DeflaterOutputStream format: RFC 1950 container around RFC 1951 DEFLATE). Since Objo has no compression API, I implemented zlib in pure Objo (a port of zlib's puff.c plus Adler-32). It is correct, but at current loop speed it is the dominant cost: 3.2 s for 7.7 KB.
What would help, ideally as shared methods on a Compression class or on MemoryBlock:
Compression.ZlibCompress(data As MemoryBlock) As MemoryBlock # RFC 1950 (78 xx header, Adler-32 trailer)
Compression.ZlibDecompress(data As MemoryBlock) As MemoryBlock
Compression.DeflateCompress / DeflateDecompress # raw RFC 1951, no header
Compression.GZipCompress / GZipDecompress # RFC 1952, for files and HTTP
An optional compression level on the compress side would be nice but isn't essential. .NET's ZLibStream / DeflateStream / GZipStream cover all three directly.
3) Request: bulk MemoryBlock operations
With 25 µs per iteration, any byte-by-byte loop is expensive, so native bulk operations would make a large difference. The ones I miss most:
mb.CopyBytes(source As MemoryBlock, sourceOffset As Integer, destOffset As Integer, count As Integer)
mb.Mid(offset As Integer, length As Integer) As MemoryBlock # or Slice
mb.ToText(offset As Integer, length As Integer) As String # UTF-8 decode of a range
mb.Resize(newSize As Integer) # keep contents
mb.IndexOf(value As Integer, startAt As Integer) As Integer
If any of these already exist under other names, a pointer would be great and I'll switch to them straight away.
Happy to send the complete project (pure Objo zlib, B4XSerializator, jRDC client) as a test case - it also makes a decent stress test for the VM. Thanks for looking into it!
Best regards,
Guido