Category Archives: Math and AI

A Story of the Determinant of Operator-Valued Semicircular Elements

Operator-valued semicircular elements are the powerhouse of free probability and random matrix theory. In particular, the intrinsic freeness principle typically replaces random matrices by operator-valued semicircular elements and shows that, in many respects, the former are close to the latter. Thus it is important to understand as much as possible about operator-valued semicirculars.

An operator-valued semicircular element – say over B=M_m(\mathbb C) – is of the form

S=a1+i=1naisiwitha,a1,,anMmsa(𝐂),S=a\otimes 1+\sum_{i=1}^n a_i\otimes s_i\qquad\text{with} \qquad a,a_1,\dots,a_n\in M_m^{sa}(\mathbf C),

where s_1,\dots,s_n are free semicirculars. We denote by {\rm tr}_m the normalized matrix trace on B, and by \tau the trace on the algebra, where the s_i live. All information about such an element S is contained in its

mean(idτ)[S]=aand its covarianceη(b)=:(idτ)[(Sa)b(Sa)]=i=1naibai.\text{mean} \quad ({\rm id} \otimes \tau)[S]=a \qquad \text{and its covariance}\quad \eta(b)=:({\rm id}\otimes \tau)[(S-a)b(S-a)]=\sum_{i=1}^n a_i b a_i.

The covariance is a completely positive map from B to B.

We also know how, in principle, to calculate the distribution \mu_S of S. Its scalar-valued Cauchy transform is

gS(z)=(trmτ)[(z1S)1].g_S(z)=(\tr_m\otimes \tau)[(z1-S)^{-1}].

This can be extracted from the operator-valued Cauchy transform

GS(b)=(idτ)[(bS)1]G_S(b)=({\rm id}\otimes\tau)[(b-S)^{-1}]

simply by

gS(z)=trm[GS(z1)].g_S(z)=\tr_m[G_S(z1)].

And G_S is determined by the nice quadratic matrix-valued equation, usually called the Dyson equation,

(ba)GS(b)=1+η(GS(b))GS(b). (b-a)\cdot G_S(b)=1+\eta(G_S(b))\cdot G_S(b).

So everything is determined by mean and covariance. Of course, some quantities are easier, others harder to extract in a meaningful way. For moments of S, for example, one has the operator-valued free version of the Wick/Isserlis formula. The operator norm of S, on the other hand, is not so easily accessible. In principle, this is given by the asymptotics of high moments, but that is not very concrete. So the following celebrated formula of Franz Lehner

λmax(a1+i=1naisi)=infbB,b>0λmax(b+a+η(b1))\lambda_{\max} \left(a\otimes 1+\sum_{i=1}^n a_i\otimes s_i\right) = \inf_{b\in B, b \gt 0}\lambda_{\max} \left(b+a+\eta(b^{-1})\right)

comes as a nice surprise, as it reduces the upper spectral edge of the infinite-dimensional S to a variational problem for the same quantity on the finite-dimensional B. Here \lambda_{\max} denotes the upper edge of the spectrum; for positive a, Lehner’s original formulation gives the corresponding operator norm.

A few years ago together with Tobias Mai I was trying to understand the determinant (more precisely, the Fuglede–Kadison determinant) of an operator-valued semicircular element, in order to get some info about the accumulation of mass of the distribution of such a semicircular element close to zero. And we succeeded in finding a nice formula in terms of the covariance; more precisely the determinant was essentially given by the capacity of the covariance map. Capacity here is the quantity introduced by Gurvits in connection with operator scaling; it should not be confused with the various notions of channel capacity in quantum information. Unwinding this notion of capacity one can write our result concretely as

Δ(i=1naisi)=e1/2infbB,b>0Δ(bη(b1))1/2.\Delta(\sum_{i=1}^n a_i\otimes s_i) = e^{-1/2}\cdot \inf_{b\in B, b>0} \Delta\left(b\cdot \eta(b^{-1})\right)^{1/2}.

\Delta is here the Fuglede–Kadison determinant, which reduces in the finite-dimensional situation to

Δ(b)=det(|b|)1/mforbMm().\Delta(b)=\det (\vert b\vert)^{1/m}\qquad\text{for}\qquad b\in M_m(\mathbb C).

As you see, our result is only for mean a=0. But in this case it has a striking similarity with Lehner’s formula:

i=1naisi=infbB,b>0b+η(b1),\| \sum_{i=1}^n a_i\otimes s_i\| = \inf_{b\in B, b>0} \left\|b+\eta(b^{-1})\right\|,

And of course it raises the question: Can we extend our formula also to the case of general mean. We thought a little bit about this way back then, but did not see how to do this.

Now, in the times of AI hype, it is tempting to come back to this question and ask ChatGPT what it has to say on this. And actually, it has something to say. It proposes the following generalization

Δ(a1+i=1naisi)=e1/2infb>0, 𝒞bη0Δ(b)exp{12a,(𝒞b+η)1(a)2},\Delta\left(a\otimes 1+\sum_{i=1}^n a_i\otimes s_i\right) = e^{-1/2} \inf_{b \gt 0,\ \mathcal C_b-\eta\succeq 0} \Delta(b)\, \exp\left\{ \frac12 \left\langle a, (\mathcal C_b+\eta)^{-1}(a) \right\rangle_2 \right\},

where

𝒞b(x)=bxb,andx,y2=trm(xy)forx,yMm()\mathcal C_b(x)=bxb,\qquad\text{and}\qquad\langle x,y\rangle_2 = \operatorname{tr}_m(x^*y)\qquad \text{for} \qquad x,y\in M_m(\mathbb C)

is the Hilbert-Schmidt inner product.

The condition \mathcal C_b-\eta\succeq 0 means positivity as an operator on the Hilbert–Schmidt space M_m(\mathbb C), i.e.

x,𝒞b(x)2x,η(x)2,or equivalentlytrm(xbxb)trm(xη(x)),for all xMm().\left\langle x,\mathcal C_b(x)\right\rangle_2 \geq \left\langle x,\eta(x)\right\rangle_2,\qquad\text{or equivalently}\qquad \operatorname{tr}_m(x^*bxb) \geq \operatorname{tr}_m\left(x^*\eta(x)\right), \qquad\text{for all }x\in M_m(\mathbb C).

ChatGPT also produces what it claims is a proof of all this — note that even the reduction of this to my formula with Tobias in the case of a=0 is non-trivial — and makes some remarks on the relevance for the Brown measure of an operator-valued circular element. Recall that the Brown measure of an operator TT is encoded by the logarithmic potential

zlogΔ(Tz1).z\mapsto \log \Delta (T-z1).

So understanding determinants of shifted operators is exactly what one needs here.

All this looks somehow reasonable, but not very enlightening; and even if I had a Lean certificate for all this (which I don’t), I would still not be too satisfied.

So let’s talk about all this and hope that, together, we can get a better understanding of what is really going on.

Doing Mathematics with AI — but How?

AI is revolutionizing mathematics, and it will not go away. So somehow we have to arrange ourselves with it and find our way of doing mathematics under these changed conditions.

I am of course also fascinated by asking AI (ChatGPT in my case) to solve all my problems. And sometimes there are indeed answers which seem to be okay and which promise some real progress. But those answers are usually complicated and quite technical, and I don’t really feel like becoming a little helper and argument checker for AI.

I agree with many voices saying that asking the meaningful questions, and understanding and presenting the answers, will remain some of the main tasks for us. But this should happen in a way that we still feel good about going along with it. For me this means that I would like to take from AI mainly some ideas, but then think myself about whether they could be true, what they actually mean, how one could prove them, and what is the best way to present them and convince others.

Of course, like everybody else, I don’t have a final answer to how this should work in practice, or whether this is a route which will help mathematics to survive in anything resembling its present form.

Anyhow, I suppose we just have to try different routes. Here is one experiment I would like to make.

At the moment I am thinking about determinants of operator-valued semicircular elements. Together with Tobias Mai, we derived a few years ago a formula for the determinant of such an element S in terms of its covariance map. This was, however, for the case of vanishing mean of S, and I wondered whether there could be a nice and useful extension to the case of general mean.

Okay, so I asked ChatGPT. And it did what two years ago would have seemed totally out of range, but has now become almost ordinary: it gave me an answer, proved it, and also made some comments on possible implications for questions about the Brown measure of operator-valued circular elements.

It even wrote a manuscript about all this.

I could now just put my name on it, say that the ideas were developed in conversations with AI, and declare that I take full responsibility for the results. But this does not feel right to me. And, more importantly, it does not feel very satisfying.

So here is what I would rather like to do.

I will give some background, explain what ChatGPT proposes as the answer, and then open the problem for public discussion: Is the statement correct? Is it perhaps obvious? How is it related to things that somebody already knows? What is the right proof? And what other interesting questions or conjectures might come out of it?

My dream would be that the answers — and also the questions — are actually self-thought, and not just copies of something another AI conversation produced. Of course everybody could now ask AI about the problem and perhaps produce a paper on it. But this should not be about publishing or priority. It should be about understanding.

Maybe one could think of it as telling a mathematical story together, in a somewhat grassroots way: somebody starts with a question, somebody else recognizes a connection, another person finds an argument, somebody points out that the whole thing is wrong, or suggests the right formulation, and gradually we understand what is really going on. Of course, we can use AI in the background for information, references, or inspiration. But AI should not tell the story.

I have no idea whether this will work.

But let us try.

In my next post I will give the concrete mathematical problem and the beginning of the story.