On Thu, Aug 20, 2026 at 06:51:08PM +0100, Matthew Wilcox wrote:
> On Thu, Aug 20, 2026 at 07:16:05PM +0200, David Hildenbrand (Arm) wrote:
> > > Consider this real deadlock pattern that lockdep cannot detect:
> > >
> > > context X context Y context Z
> > >
> > > mutex_lock A
> > > folio_lock B
> > > folio_lock B <- DEADLOCK
> > > mutex_lock A <- DEADLOCK
> > > folio_unlock B
> > > folio_unlock B
> > > mutex_unlock A
> > > mutex_unlock A
> >
> > But that really just boils down to folio lock being implemented as a PG_lock +
> > some advanced wait mechanism. And we must do that because of lack of bits in
> > struct page.
> >
> > Willy mentioned in a previous version [1]: "I don't think it makes sense to
> > track lock state in the page (nor folio). Partly because there's just so many
> > of them, but also because the locking rules don't really apply to individual
> > folios so much as they do to the mappings (or anon_vmas) that contain folios."
> >
> > Given that lockdep is a debug feature, and we will at some point allocate struct
> > folio separately, I assume we could just squeeze a "struct lockdep_map" in there
> > in such debug configs and the world would not collapse.
> >
> > Doing that today (one "struct lockdep_map" in each "struct page") wouldn't work
> > as mm_zero_struct_page() would not expect such large "struct page". But
> > conceptually, for a debug kernel with a special CONFIG_LOCKDEP_PAGE_LOCK, maybe
> > that would already be ok and we could just do that (and optimize it as we
> > allocate folios separately).
> >
> > Not that it's ideal, but for a debug feature to at least check PG_lock, probably
> > an easier way to achieve it than some completely new infrastructure.
> >
> > Now, Willy said "locking rules don't really apply to individual folios", I
> > wonder if that could just help to also let lockdep check PG_lock with less
> > metadata? (didn't fully wrap my head around the implications)
> >
> > [1]
> > https://lore.kernel.org/all/aR3WHf9QZ_dizNun@casper.infradead.org/?utm_sour…
>
> There are a few things going on that make PG_lock special. Let me try
> to explain again, only better this time.
>
> 1. The current lifetime of a struct page is the lifetime of the system.
> But the semantics of its PG_lock bit change each time it is freed and
> allocated.
Yes, it's a classification issue that is very important.
> 2. The position of PG_lock in the locking hierarchy only depend on
> what the folio is currently being used for. That is, all folios in
> a given xfs inode behave exactly the same from a locking perspective.
You are exactly explaining what the classification means. Perfect.
> There's no need to build up state about how each PG_lock is used;
> they can all share. Arguably all xfs file inodes are the same as
Right. That's why DEPT doesn't use a full map in each page but just
uses a timestamp in each. For the classification, DEPT uses a few
classes for folio, using global maps:
1. folios in mm paths
2. folios in block device buffer (meta data)
3. folios in regular file cache
However, yes. I bet you could be a big help when classifying folios
more presicely according to its usage. But the current classification
is still a good start I think.
> each other (directory inodes might be different from file inodes),
> so we might want to go further than telling DEPT that "this folio
> belongs to this inode" and go to "this folio belongs to this xfs file
> inode".
Totally agree.
> 3. PG_lock can be taken in task context then released in interrupt
> context. For full points, we need to mark the exact point at which
> we submit the folio for read. Otherwise we can get into the situation
> alluded to by f2c817bed58d and better discussed at
> https://lore.kernel.org/linux-mm/20200127150024.GN1183@dhcp22.suse.cz/
> where we have the folio locked but haven't yet submitted it for I/O
> so it doesn't matter how long we wait, it will never come unlocked.
Interesting.
The following abstraction might make DEPT work with it. For example:
Annotate the point submitting IO as an event for the folio_lock() to
be released. That way, the issue above can be detected by DEPT.
Again, DEPT can do every thing we need w.r.t. deadlock.
Byungchul
On Thu, Aug 20, 2026 at 07:16:05PM +0200, David Hildenbrand (Arm) wrote:
> On 7/6/26 08:18, Byungchul Park wrote:
> > Hi Linus and folks,
>
> Hi,
Hi,
> I think there was plenty of feedback from locking maintainers in the past. One
> question and a comment below.
>
> >
> > DEPT(DEPendency Tracker) is a runtime deadlock detection framework that
> > sees what lockdep cannot.
> >
> > I'm thrilled to share that DEPT has moved beyond theory and is now
> > catching real deadlocks in the wild:
> >
> > https://lore.kernel.org/lkml/6383cde5-cf4b-facf-6e07-1378a485657d@I-love.SA…
> > https://lore.kernel.org/lkml/1674268856-31807-1-git-send-email-byungchul.pa…
> > https://lore.kernel.org/all/b6e00e77-4a8c-4e05-ab79-266bf05fcc2d@igalia.com/
> >
> > I've added comprehensive documentation explaining DEPT's design and usage.
> > Getting started is as simple as enabling CONFIG_DEPT and watching dmesg.
> >
> > THE PROBLEM LOCKDEP CANNOT SOLVE
> > --------------------------------
> >
> > Lockdep has been our trusted deadlock detector for two decades, but it
> > has a fundamental blind spot: it tracks lock acquisition order, not the
> > actual waits and events that cause deadlocks. This means lockdep misses:
> >
> > * Deadlocks involving folio locks (not released within the context)
> > * Cross-context synchronization like wait_for_completion()/complete()
> > * DMA fence waits, RCU waits, and general waitqueue patterns
> > * Any synchronization primitive outside the classic lock/unlock model
> >
> > Consider this real deadlock pattern that lockdep cannot detect:
> >
> > context X context Y context Z
> >
> > mutex_lock A
> > folio_lock B
> > folio_lock B <- DEADLOCK
> > mutex_lock A <- DEADLOCK
> > folio_unlock B
> > folio_unlock B
> > mutex_unlock A
> > mutex_unlock A
>
> But that really just boils down to folio lock being implemented as a PG_lock +
> some advanced wait mechanism. And we must do that because of lack of bits in
> struct page.
>
> Willy mentioned in a previous version [1]: "I don't think it makes sense to
> track lock state in the page (nor folio). Partly because there's just so many
> of them, but also because the locking rules don't really apply to individual
> folios so much as they do to the mappings (or anon_vmas) that contain folios."
Exactly. That's why we use classification e.g. lock class - DEPT also
makes use of the concept.
DEPT doesn't use a full map in each page but uses a minimum space for a
timestamp in each to track when each starts to wait so as to use the
recorded timestamp when the event occurs e.g. folio_unlock().
> Given that lockdep is a debug feature, and we will at some point allocate struct
> folio separately, I assume we could just squeeze a "struct lockdep_map" in there
> in such debug configs and the world would not collapse.
That's a good news for lockdep. (And even for DEPT :)
> Doing that today (one "struct lockdep_map" in each "struct page") wouldn't work
> as mm_zero_struct_page() would not expect such large "struct page". But
> conceptually, for a debug kernel with a special CONFIG_LOCKDEP_PAGE_LOCK, maybe
> that would already be ok and we could just do that (and optimize it as we
> allocate folios separately).
Sounds great.
> Not that it's ideal, but for a debug feature to at least check PG_lock, probably
> an easier way to achieve it than some completely new infrastructure.
I understand what you are going to tell.
However, it's worth noting that lockdep tracks dependencies basically
based on **lock acqusition orders** in the system. To make it track
even rwlock and general synchronization mechanism as well, lockdep has
no choice but to get more complicated.
Focusing on only the dependency checking, the most parts of lockdep are
for the tricky things, so the reusable parts are not that big.
> Now, Willy said "locking rules don't really apply to individual folios", I
> wonder if that could just help to also let lockdep check PG_lock with less
> metadata? (didn't fully wrap my head around the implications)
That's what DEPT did and what brought external wgen introduced in DEPT.
I was considering the exactly same thing :)
Again, lockdep that tracks lock acquisition orders can't do that.
> [1]
> https://lore.kernel.org/all/aR3WHf9QZ_dizNun@casper.infradead.org/?utm_sour…
>
>
> It's your guiding example, that's why I mention it. You do mention other wait
> cases here, I don't know anything about them, but for folios it's really just
> "we used a single bit so far" AFAIKs.
It doesn't matter whether it's implemented using bit or not. folio lock
is quite special since it's allowed to be released other than the
acquisition context that makes lockdep impossible to track them.
> [...]
>
> >
> > Q. Why not build DEPT into lockdep?
> >
> > A. Lockdep is stable, battle-tested code. I chose separation because
> > while DEPT borrows BFS and hashing ideas, the wait/event model
> > requires rebuilding from scratch. Lockdep was designed for lock
> > acquisition order — retrofitting it would risk its stability.
>
> Why can't this just be some configurable extension to lockdep
> (CONFIG_LOCKDEP_XYZ) until the feature is stable and can unconditionally be
> enabled along with it?
Answered?
> I don't quite buy the "would risk its stability" argument. A lot of stuff we do
> "risks stability", every day :)
That's awsome anyway :)
> Is there another good reason (incompatible with X, dangerous with Y, cinfusing
> Z) why this really must be a separate thing?
Roughly:
1. Similar or less effort is needed for the new one - retrofitting
lockdep is not easy and big changes are required since the
reusable parts are not that big.
2. Even though you didn't agree, retrofitting it would risk its
stability.
> >
> > Q. Will DEPT replace lockdep?
> >
> > A. No. Lockdep validates correct lock usage — that's not going away.
> > DEPT supersedes only the dependency-checking logic when mature.
>
> It's quite unfortunate that we'd end up with another similar-but-different
> mechanism, that will just end up confusing people.
I meant, at least dependency checking engine should be altered, but you
make sense. Worth thinking it more.
> But I am not a locking maintainer. I think there was plenty of discussion in the
> past, so I might just be raising points that were already discussed in the past,
> but I really just read some random pieces of earlier discussions. (ideally
> previous discussions would be summarized here)
>
> Long story short: we are now in v19 and I think there was pushback in the past.
> Did the opinion of locking maintainers change, or is there a way forward to
> integrate this in a way that would make locking maintainers accept this?
One of locking maintainers who I met in an LPC told me that he agrees
with the direction of DEPT and supports DEPT, not officially tho.
What he and other people are concerning w.r.t DEPT the most is, false
positives, which is the most important issue for now.
At the same time, I think the most important thing is to make DEPT
useful in practice especially with folio locks involved. Actually, I'm
planning to share DEPT's true reports periodically to LKML and work with
people who believe DEPT can make things better.
Any advices will be welcome. Thanks for your opinions.
Byungchul
> --
> Cheers,
>
> David
set_memory_decrypted() doesn't (currently) guarantee to preserve or zero
memory, but both the GICv3 ITS driver and the system_cc_shared dma-buf
heap currently allocate memory with __GFP_ZERO followed by calling
set_memory_decrypted(). On an Arm CCA system with MEC this can cause
ciphertext to be visible to the guest rather than the expected zeros.
Patches 1 and 3 fix this by zeroing after the set_memory_decrypted()
call.
Patches 2 and 4 fix other related bugs that Sashiko found. Patch 2 fixes
the issue that set_memory_decrypted() can be a sleeping call, so moves
the allocation out of an atomic context.
Patch 4 deals with the situation where set_memory_decrypted() fails and
the rollback path could attempt to re-encrypt memory which was never
decrypted.
I've sorted the patches by area, but there's no (semantic) dependency
between them.
Changes in v2:
* Switched to use BIT(n) rather than 1 << n in the GICv3 change.
* Added Jason's R-b.
* Patches 2 and 4 are new.
v1: https://lore.kernel.org/r/20260820105026.53208-1-steven.price@arm.com
Steven Price (4):
irqchip/gic-v3-its: Zero shared pages after conversion
irqchip/gic-v3-its: Allocate VPE tables from sleepable context
dma-buf: heaps: Zero system shared heap pages after conversion
dma-buf: heaps: Fix shared system heap allocation rollback
drivers/dma-buf/heaps/system_heap.c | 24 +++++++++++++-----
drivers/irqchip/irq-gic-v3-its.c | 39 ++++++++++++++++-------------
2 files changed, 40 insertions(+), 23 deletions(-)
--
2.43.0
Arm CCA includes "Memory Encryption Contexts" (MEC) which allows the
private and shared data accessible to a guest to have different memory
encryption keys. Consequently when converting memory to shared, the
memory encryption key used to access the physical page will change.
Both the GICv3 ITS driver and the system_cc_shared dma-buf heap
currently allocate memory with __GFP_ZERO and then decrypt it. With MEC
the zeroing is done with the wrong encryption key and the data visible
after decryption may be ciphertext. The RMM is required to scrub the
data, but may perform this scrub with a different encryption key to the
eventual key that will be used for shared access.
Fix these two sites by avoiding the __GFP_ZERO during the allocation and
performing a clear_pages() call after the decryption.
Steven Price (2):
irqchip/gic-v3-its: Zero shared pages after conversion
dma-buf: heaps: Zero system shared heap pages after conversion
drivers/dma-buf/heaps/system_heap.c | 14 +++++++++++---
drivers/irqchip/irq-gic-v3-its.c | 7 ++++++-
2 files changed, 17 insertions(+), 4 deletions(-)
--
2.43.0
On 20/08/2026 12:07, Dmitry Baryshkov wrote:
> On Thu, Aug 20, 2026 at 11:07:45AM +0200, Krzysztof Kozlowski wrote:
>>>>>>
>>>>>> Device node with this compatible is already populated, so this looks
>>>>>> simply wrong or you are adding a duplicated driver.
>>>>>>
>>>>>> That's a no-go, you are supposed to work with existing drivers and grow
>>>>>> them.
>>>>> I'll bring the discussion again here, there was a discussion to move the
>>>>> driver to accel subsystem if we want to support new features/uAPI
>>>>> changes. Please read [1],[2] threads. The intention is to replace
>>>>> fastrpc driver with QDA eventually.
>>>>
>>>> None of them address the problem. You want to grow fastrpc into user of
>>>> dmabuf? So you move it from misc to here.
>>>
>>> It's not as easy and nice, so I think in this case it's better to repeat
>>
>> I disagree. The existing fastrpc driver is not that complicated. It's
>> actually moderate amount of code, much less than Venus was (~7 times less).
>>
>> It easily can grow to support two interfaces and the only difficulty is
>> how to manage these two interfaces simultaneously or exclusively, e.g.
>> opening first one disables the second.
>
> I see the point here.
>
> Would it be acceptable if we add QDA support only on the new platforms
> (e.g. via the SoC-specific compat), provide QDA for those platforms,
> and, once it reaches complete API and feature parity, we remove the old
> fastrpc driver, migrati old platforms.
The problem with this approach is that we have no guarantees that it
will reach feature parity in respect of old interface, thus old driver
might stay forever. If we agree for duplicated driver, contributors have
no incentives to support old approach.
Much better is to refine the old driver, gradually adding new features
while maintaining old stuff. This is the only way we can force
contributors to actively work on minimizing duplicate parts.
Best regards,
Krzysztof
On Thu, Aug 20, 2026 at 01:32:05PM +0100, Marc Zyngier wrote:
> On Thu, 20 Aug 2026 11:50:24 +0100,
> Steven Price <steven.price(a)arm.com> wrote:
> >
> > its_alloc_pages_node() passes __GFP_ZERO to the page allocator before
> > calling set_memory_decrypted(). This assumes that converting a page from
> > private to shared preserves its contents.
> >
> > For Arm CCA with MEC (Memory Encryption Contexts) the key used to access
> > the page will change, and so by default the visible data will change.
> > The host could ensure that it zeros the page, but rather than relying on
> > the host's behaviour it's best if the guest simply zeros after the
> > decryption rather than before. Specifically in this case the ITS tables
> > are required to be zeroed.
>
> What are the guarantees that we want to enforce post decryption? My
> recollection is that the RME firmware cleans the caches to the PoPA,
> making the data immediately visible to the hypervisor. Obviously, this
> isn't the case anymore, since the zeroing comes after that, and I
> don't see any CMO enforcing this.
I thought any CMO stuff was principally about cleaning things as part
of the MEC change? Coherency after the memory is made shared should
follow the normal cachable memory model rules, just like in a non-CC
VM? We don't need further explicit CMOs for that.
Post decryption I would expect from all architectures:
1) Neither the guest or host take a fault/error when accessing the
memory. ie the host may immediately pass this memory to an
O_DIRECT system call and have its kernel read from it.
It must not crash the kernel.
2) So long as the memory is mapped cachable it should follow the
normal memory model visibility rules. ie it works the same as
VM CPU memory prior to CC
3) Rules for actual DMA are the same as prior to CC, the VM is
expected to issue its own flushes prior to DMA if the platform
requires it.
Given the requirements for #1, is there actually any case on any
platform where the host doesn't *have* to fill the memory? Is there a
platform with MEC that doesn't generate an error on reading with the
wrong MEC? Without MEC it surely has to be zero'd in the RMM world,
right?
I've argued before that set_memory_decrypted() should be defined to
return 0'd memory. I think there are real systems that *have* to zero
the memory as part of the state change and this API is now
forcing an extra zeroing.
> I'm concerned that this relies on undocumented behaviours that may
> hold today on some undisclosed combinations of HW and hypervisors, but
> that are not guaranteed at all. set_memory_decrypted() doesn't really
> say anything, and I have the feeling that we may want some hypervisor
> specific hook to perform the correct CMO magic. I don't think this is
> required right now, but I'm not excluding anything!
I would expect any required CMOs to be part of the arch's
implementation of set_memory_decrypted()?
Jason