A while ago, I spent quite a bit of time in the parts of our product that read metadata and MSIL bytecode. By pure chance, I stumbled across this MSIL instruction definition:

public static readonly Opcode No = new Opcode("no.",
  StackBehavior.Pop0, StackBehavior.Push0,
  OperandType.SByte, OpcodeType.Prefix, 2,
  OpcodeValue.No, FlowControl.Meta, false, 0);

I’d never seen this opcode before. The definition makes it clear that it’s a prefix rather than a standalone instruction. The only other prefixes that came to mind were tail. and constrained., though there are others. What made it even more intriguing was that the .NET Framework’s System.Reflection.Emit.OpCodes class doesn’t define it at all. Even the detailed and generally excellent book Expert .NET 2.0 IL Assembler doesn’t mention it once. But Mono.Cecil does have this prefix, so the people writing metadata readers and writers clearly knew something the rest of us didn’t. And it wasn’t just the mysterious name, no., that caught my attention: it also takes an sbyte operand. You can not only attach this mysterious prefix to an instruction, but somehow control what it does.

My first attempt at Googling it didn’t turn up anything useful either. Oddly enough, it was the CLI standard that finally revealed the secret of opcode 0xFE19:

The CLI specification for the no. prefix, which allows type, range, and null checks to be skipped.

Now it all made sense. This opcode was supposed to let you disable the checks that virtually every CLR application lives with: automatic null checks, array bounds checks, and type checks when downcasting or storing into those dreadful covariant arrays borrowed from the JVM. These are, after all, some of the very checks that make a managed runtime “managed.”

Of course, bytecode with a no. prefix is no longer “verifiable,” but who cares? At that point, the real world ceased to exist for me. Ten minutes later, I had a Mono.Cecil rewriter consisting of a single loop, inserting the no. prefix everywhere it made sense…

The test application crashes, while ILDasm displays the no. prefix bytes as an unknown instruction.

An ExecutionEngineException brought me back to reality and killed the excitement. ILDasm refused to display opcode 0xFE19, even though the byte stream should have represented a valid no. prefix with an operand. I suspected Mono.Cecil was emitting something incorrectly, so I tried writing no. by hand and assembling it with ILAsm. Unknown token, it said.

Later, I found an old thread on the MSDN forums that confirmed my suspicions: despite being part of the CLI standard, this prefix had simply never been implemented in either Rotor or the CLR…

That was thoroughly disappointing, but it made me want to write this post and daydream about what we could do with no.. Sure, it lets you bypass managed safety guarantees, but it could also eliminate a substantial part of their overhead. Adding an item to a List<T> could just store it in the backing array with a single bounds check, instead of doing two: one in the list’s implementation and another implicit check on the array store. On top of that, arrays of reference types incur a pointless type check, even though List<T> itself is invariant. The examples of redundant checks are endless. Yes, all these checks add a uniform layer of O(1) overhead to your code, and in 99% of cases they aren’t a serious problem. But why shouldn’t we be able to get rid of them when we want to?

Obviously, a JIT compiler has a very tight time budget and can’t afford many optimizations, especially one without an interpreter, profiling, and the million other things you find in runtimes like HotSpot. But the CLR takes this to an extreme: even a simple loop filling an int[] with the number 42 takes twice as long going from the end to the beginning as it does going from the beginning to the end, all because of bounds checks. What’s more, JIT compilation will never be able to perform certain optimizations that require extensive static analysis. And that’s fine: JIT compilation shines at optimizations based on profiling or knowledge of a particular platform. But we shouldn’t stop there. Many transformations can be applied to the bytecode itself, and we could also leave hints for the JIT compiler to make its job easier.

This raises some deeper questions. Who should define what “verifiable” means? Is verifiable, managed code really that important, especially on mobile platforms? What if verification were decoupled from the virtual machine? How often is the code running in a VM fully verifiable anyway?

A real-world example: our product uses P/Invoke occasionally, for things like the Win32 API and LevelDB. That means unmanaged code, once it gets hold of a pointer into the managed heap, can wreak as much havoc as it likes. Does that worry us? Not at all. Does it cause problems in practice? No: the unmanaged code behaves itself and doesn’t do anything nasty. Could it wreak havoc? Of course. Can we call the entire product verifiable? Of course not.

It’s a real shame that the CLR never implemented the no. prefix. It could have encouraged alternative verification tools and smarter compiler backends that use no. to eliminate managed checks proven to be redundant. Some checks can already be removed using the unmanaged operations available today. For example, you can use ldind (a pointer dereference) instead of ldelem (an array element access, essentially the same dereference plus a bounds check). But these techniques mostly work with unmanaged types: non-generic value types whose fields are also of unmanaged types.

Another bit of managed runtime behavior we could control through bytecode hints is allocation. We could allow newobj to allocate managed objects on the stack, or teach initobj to accept reference-type tokens. Then we could perform escape analysis before JIT compilation and reduce unnecessary pressure on the garbage collector, again giving up verifiability in the traditional CLI sense. To allocate unmanaged classes on the stack today, C++/CLI, for example, uses initobj and value types, which are then passed down the call stack by value or by unmanaged reference. But this approach comes with all the limitations of value types: unmanaged C++/CLI classes can’t implement managed interfaces, because that would require boxing. So the CLR currently has no proper way to allocate a managed object, complete with a real object header, on the stack.

I’d love to see a modern VM for managed languages that made it easy to control these managed runtime guarantees through bytecode. What do you think?