kafka

Commit Graph

Author	SHA1	Message	Date
Andrew Schofield	86baac103b	MINOR: Improve client error messages for share groups not enabled (#19688 ) CI / build (push) Waiting to run Details As mentioned in https://github.com/apache/kafka/pull/19378#pullrequestreview-2775598123, the error messages for a 4.1 share consumer could be clearer for the different cases for when it cannot successfully join a share group. This PR uses different error messages for the different cases such as out-of-date cluster or share groups just not enabled. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-05-13 10:42:40 +01:00
Bolin Lin	6eafe407bd	MINOR: Fix unchecked type warnings in several test classes (#19679 ) * In ConsoleShareConsumerTest, add `@SuppressWarnings("unchecked")` annotation in method shouldUpgradeDeliveryCount * In ListConsumerGroupOffsetsHandlerTest, add generic parameters to HashSet constructors * In TopicsImageTest, add explicit generic type to Collections.EMPTY_MAP to fix raw type usage Reviewers: Ken Huang <s7133700@gmail.com>, TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-13 14:59:22 +08:00
ChickenchickenLove	62bec20aef	KAFKA-19242: Fix commit bugs caused by race condition during rebalancing. (#19631 ) ### Motivation While investigating “events skipped in group rebalancing” ([spring‑projects/spring‑kafka#3703](https://github.com/spring-projects/spring-kafka/issues/3703)) I discovered a race condition between - the main poll/commit thread, and - the consumer‑coordinator heartbeat thread. If the main thread enters `ConsumerCoordinator.sendOffsetCommitRequest()` while the heartbeat thread is finishing a rebalance (`SyncGroupResponseHandler.handle()`), the group state transitions in the following order: ``` COMPLETING_REBALANCE → (race window) → STABLE ``` Because we read the state twice without a lock: 1. `generationIfStable()` returns `null` (state still `COMPLETING_REBALANCE`), 2. the heartbeat thread flips the state to `STABLE`, 3. the main thread re‑checks with `rebalanceInProgress()` and wrongly decides that a rebalance is still active, 4. a spurious `CommitFailedException` is returned even though the commit could succeed. For more details, please refer to sequence diagram below. <img width="1494" alt="image" src="https://github.com/user-attachments/assets/90f19af5-5e2d-4566-aece-ef764df2d89c" /> ### Impact - The exception is semantically wrong: the consumer is in a stable group, but reports failure. - Frameworks and applications that rely on the semantics of `CommitFailedException` and `RetryableCommitException` (for example `Spring Kafka`) take the wrong code path, which can ultimately skip the events and break “at‑most‑once” guarantees. ### Fix We enlarge the synchronized block in `ConsumerCoordinator.sendOffsetCommitRequest()` so that the consumer group state is examined atomically with respect to the heartbeat thread: ### Jira https://issues.apache.org/jira/browse/KAFKA-19242 https: //github.com/spring-projects/spring-kafka/issues/3703 Signed-off-by: chickenchickenlove <ojt90902@naver.com> Reviewers: David Jacot <david.jacot@gmail.com>	2025-05-12 17:01:29 +02:00
Matthias J. Sax	b66729e231	MINOR: fit HTML markup (#19676 ) CI / build (push) Waiting to run Details Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-05-11 16:20:25 -07:00
xijiu	3696c49788	KAFKA-19220 Add tests to ensure the internal configs don't return by public APIs by default (#19650 ) Add tests to check whether the results returned by the API `createTopics` and `describeConfigs` contain internal configurations. Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, TengYao Chi <frankvicky@apache.org>, TaiJuWu <tjwu1217@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-10 23:13:58 +08:00
Stanislav Kozlovski	0bc8d0c962	MINOR: Add documentation about KIP-405 remote reads serving just one partition per FetchRequest (#19336 ) [As discussed in the mailing list](https://lists.apache.org/thread/m03mpkm93737kk6d1nd6fbv9wdgsrhv9), the broker only fetches remote data for ONE partition in a given FetchRequest. In other words, if a consumer sends a FetchRequest requesting 50 topic-partitions, and each partition's requested offset is not stored locally - the broker will fetch and respond with just one partition's worth of data from the remote store, and the rest will be empty. Given our defaults for total fetch response is 50 MiB and per partition is 1 MiB, this can limit throughput. This patch documents the behavior in 3 configs - `fetch.max.bytes`, `max.partition.fetch.bytes` and `remote.fetch.max.wait.ms` Reviewers: Luke Chen <showuon@gmail.com>, Kamal Chandraprakash <kamal.chandraprakash@gmail.com>, Satish Duggana <satishd@apache.org>	2025-05-10 16:48:55 +05:30
Alyssa Huang	042be5b9ac	MINOR: Fix some Request toString methods (#19655 ) CI / build (push) Waiting to run Details Reviewers: Colin P. McCabe <cmccabe@apache.org>	2025-05-09 23:42:34 -07:00
Shivsundar R	58c08441d1	KAFKA-19229: Ignore background errors while closing share consumers. (Fix flaky test) (#19647 ) CI / build (push) Waiting to run Details - A couple of newly added tests were found to be flaky in `AuthorizerIntegrationTest.scala`. - `testShareGroupDescribeWithGroupDescribeAndTopicDescribeAcl` and `testShareGroupDescribeWithoutGroupDescribeAcl`. These tests pass locally, so could not replicate the failure. - But logs from develocity indicated that the test fails when the following condition happens : When the background error event arrives after the consumer had unsubscribed, then these events are processed in the `handleCompletedAcknowledgements` method and the exception from the event is thrown, preventing `close()` to complete. - We need to handle this race condition where we might get the background event after unsubscribe and before processing the callbacks. - PR fixes this by ignoring the exceptions in the background queue when the `handleCompletedAcknowledgements` method is called during `close()`. This ensures `close()` completes successfully. - Have added a unit test which mimics the race condition as well. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-05-09 11:20:09 +01:00
Andrew Schofield	70c0aca4b7	KAFKA-17897: Deprecate Admin.listConsumerGroups [2/N] (#19508 ) CI / build (push) Waiting to run Details Admin.listConsumerGroups() was able to use the early versions of ListGroups RPC with the version used dependent upon the filters the user specified. Admin.listGroups(ListGroupsOptions.forConsumerGroups()) inadvertently required ListGroups v5 because it always set a types filter. This patch handles the UnsupportedVersionException and winds back the complexity of the request unless the user has specified filters which demand a higher version. It also adds ListGroupsOptions.forShareGroups() and forStreamsGroups(). The usability of Admin.listGroups() is much improved as a result. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, PoAn Yang <payang@apache.org>	2025-05-09 08:38:16 +01:00
ShihYuan Lin	1ccaddaa70	KAFKA-19209: Clarify index.interval.bytes impact on offset and time index (#19657 ) Update docs to note index.interval.bytes sets entry frequency for offset index and, conditionally, time index. Improve clarity and readability of index.interval.bytes description. Reviewers: Luke Chen <showuon@gmail.com>	2025-05-09 09:48:55 +08:00
David Jacot	98e535b524	MINOR: Simplify OffsetFetchResponse (#19642 ) While working on https://github.com/apache/kafka/pull/19515, I came to the conclusion that the OffsetFetchResponse is quite messy and overall too complicated. This patch rationalize the constructors. OffsetFetchResponse has a single constructor accepting the OffsetFetchResponseData. A builder is introduced to handle the down conversion. This will also simplify adding the topic ids. All the changes are mechanical, replacing data structures by others. Reviewers: Lianet Magrans <lmagrans@confluent.io>	2025-05-08 14:57:45 +02:00
Apoorv Mittal	2dd6126b5d	KAFKA-18855 Slice API for MemoryRecords (#19581 ) CI / build (push) Waiting to run Details The PR adds `slice` API in `Records.java` and further implementation in `MemoryRecords`. With the addition of ShareFetch and it's support to read from TieredStorage, where ShareFetch might acquire subset of fetch batches and TieredStorage emits MemoryRecords, hence a slice API is needed for MemoryRecords as well to limit the bytes transferred (if subset batches are acquired). MemoryRecords are sliced using `duplicate` and `slice` API of ByteBuffer, which are backed by the original buffer itself hence no-copy is created rather position, limit and offset are changed as per the new position and length. Reviewers: Andrew Schofield <aschofield@confluent.io>, Jun Rao <junrao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-08 14:02:25 +08:00
Lianet Magrans	67b46fec15	MINOR: introduce structure to keep member assignment with topic Ids (#19645 ) - Add new DS to wrap the member assignment (containing topic Ids, names and partitions), to easily access the data as needed. This will be used in following PR to integrate assignment with topic IDs into the subscription state. - Improve logging on the client assignment/reconciliation path No changes in logic. Reviewers: TengYao Chi <frankvicky@apache.org>, Andrew Schofield <aschofield@confluent.io>	2025-05-07 13:57:56 -04:00
Kirk True	d3707fc815	KAFKA-19214: Clean up use of Optionals in RequestManagers.entries() (#19609 ) Change: `public List<Optional<? extends RequestManager>> entries();` to: `public List<RequestManager> entries();` and clean up the callers. Reviewers: TengYao Chi <kitingiao@gmail.com>, Andrew Schofield <aschofield@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-07 17:18:12 +01:00
yunchi	d034268312	MINOR: Remove ConstantBrokerOrActiveKController (#19654 ) `ConstantBrokerOrActiveKController` was introduced in #14399, to provide a mechanism for selecting the least loaded broker or the active controller when using `bootstrap.controllers`. Usage was removed in #18002, after `alterConfigs` was deprecated in Kafka 2.4.0. Reviewers: PoAn Yang <payang@apache.org>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, Ken Huang <s7133700@gmail.com>, TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-07 20:23:29 +08:00
Lan Ding	e1da318722	MINOR: add boundary IT for delivery count (#19649 ) CI / build (push) Waiting to run Details see https://github.com/apache/kafka/pull/19430#pullrequestreview-2809619176 Add boundary IT for delivery count. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-05-06 22:05:02 +01:00
Andrew Schofield	7d027a4d83	KAFKA-19218: Add missing leader epoch to share group state summary response (#19602 ) CI / build (push) Waiting to run Details When the persister is responding to a read share-group state summary request, it has no way of including the leader epoch in its response, even though it has the information to hand. This means that the leader epoch information is not initialised in the admin client operation to list share group offsets, and this then means that the information cannot be displayed in kafka-share-groups.sh. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>, Sushant Mahajan <smahajan@confluent.io>	2025-05-06 14:53:12 +01:00
Dmitry Werner	0810650da1	MINOR: Small cleanups in clients tests (#19634 ) - Removed unused fields and methods in clients tests - Fixed IDEA code inspection warnings Reviewers: Ken Huang <s7133700@gmail.com>, PoAn Yang <payang@apache.org>, Andrew Schofield <aschofield@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>, TengYao Chi <frankvicky@apache.org>	2025-05-06 20:19:21 +08:00
yunchi	4e77466f6a	KAFKA-19170 Move MetricsDuringTopicCreationDeletionTest to client-integration-tests module (#19528 ) rewrite `MetricsDuringTopicCreationDeletionTest` to `ClusterTest` infra and move it to clients-integration-tests module. Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-06 19:57:16 +08:00
Alieh Saeedi	54b3b3debc	MINOR: Convert streams group options to consumer group options in Admin APIs (#19583 ) This PR is fixing the issue introduced in #19120 The input `StreamsGroup`-options must not be ignored, but it must be converted to `ConsumerGroup`-options. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-05-06 13:26:56 +02:00
Andrew Schofield	d2bd68d50c	MINOR: Improve output for delete-offset of kafka-consumer-groups.sh (#19610 ) The output from the delete-offsets option of kafka-consumer-groups.sh can be improved. For example, the column widths are excessive which looks untidy, and the output messages can be improved. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-05-06 12:20:36 +01:00
TaiJuWu	19530738c4	KAFKA-19240 Move MetadataVersionIntegrationTest to clients-integration-tests module (#19641 ) The PR do following: 1. Move MetadataVersionIntegrationTest to clients-integration-tests module 2. rewrite to java from scala Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-05-06 00:12:57 +08:00
Shivsundar R	fedbb90c12	KAFKA-19232: Handle Share session limit reached exception in clients. (#19619 ) CI / build (push) Waiting to run Details Handle the new `ShareSessionLimitReachedException` in `ShareSessionHandler` in the client to reset the ShareSession. Added a unit test verifying the change. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-05-04 19:59:40 +01:00
yunchi	bff5ba4ad9	MINOR: replace .stream().forEach() with .forEach() (#19626 ) CI / build (push) Waiting to run Details replace all applicable `.stream().forEach()` in codebase with just `.forEach()`. Reviewers: TengYao Chi <kitingiao@gmail.com>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-04 20:39:55 +08:00
Ken Huang	c85e09f7a5	KAFKA-19060 Documented null edge cases in the Clients API JavaDoc (#19393 ) Some client APIs may return `null` values in the map, but this behavior isn’t documented in the JavaDoc. We should update the JavaDoc to include these edge cases. Reviewers: Kirk True <kirk@kirktrue.pro>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-04 20:35:02 +08:00
xijiu	b5cceb43e5	KAFKA-19205: inconsistent result of beginningOffsets/endoffset between classic and async consumer with 0 timeout (#19578 ) CI / build (push) Waiting to run Details In the return results of the methods beginningOffsets and endOffset, if timeout == 0, then an empty Map should be returned uniformly instead of in the form of <TopicPartition, null> Reviewers: Ken Huang <s7133700@gmail.com>, PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>, Lianet Magrans <lmagrans@confluent.io>	2025-05-03 13:12:20 -04:00
TengYao Chi	93e65c4539	KAFKA-18267 Add unit tests for CloseOptions (#19571 ) There is some redundant code that could be removed in `CloseOptions`. This patch also adds unit tests for CloseOptions. Reviewers: Ken Huang <s7133700@gmail.com>, PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-03 22:36:43 +08:00
Matthias J. Sax	44025d8116	MINOR: fix bug in MockConsumer (#19627 ) The setter of `maxPollRecords` wrongly checks the field instead of the argument. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, TengYao Chi <frankvicky@apache.org>	2025-05-03 14:08:18 +08:00
Sushant Mahajan	e68781414e	KAFKA-19204: Allow persister retry of initializing topics. (#19603 ) CI / build (push) Waiting to run Details * Currently in the share group heartbeat flow, if we see a TP subscribed for the first time, we move that TP to initializing state in GC and let the GC send a persister request to share group initialize the aforementioned TP. * However, if the coordinator runtime request for share group heartbeat times out (maybe due to restarting/bad broker), the future completes exceptionally resulting in persiter request to not be sent. * Now, we are in a bad state since the TP is in initializing state in GC but not persister initialized. Future heartbeats for the same share partitions will also not help since we do not allow retrying persister request for initializing TPs. * This PR remedies the situation by allowing the same. * A temporary fix to increase offset commit timeouts in system tests was added to fix the issue. In this PR, we revert that change as well. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-05-02 14:25:29 +01:00
Matthias J. Sax	f69337b37c	MINOR: use `isEmpty()` to avoid compiler warning (#19616 ) Reviewers: Anna Sophie Blee-Goldman <ableegoldman@apache.org>	2025-05-01 23:51:36 -07:00
Calvin Liu	0c1fbf3aeb	KAFKA-19073 add transactional ID pattern filter to ListTransactions (#19355 ) Propose adding a new filter TransactionalIdPattern. This transaction ID pattern filter works as AND with the other transaction filters. Also, it is empowered with Re2j. KIP: https://cwiki.apache.org/confluence/x/4gm9F Reviewers: Justine Olshan <jolshan@confluent.io>, Ken Huang <s7133700@gmail.com>, Kuan-Po Tseng <brandboat@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-05-02 00:52:21 +08:00
Lan Ding	8dbf56e4b5	KAFKA-17541:[1/2] Improve handling of delivery count (#19430 ) For records which are automatically released as a result of closing a share session normally, the delivery count should not be incremented. These records were fetched but they were not actually delivered to the client since the disposition of the delivery records is carried in the ShareAcknowledge which closes the share session. Any remaining records were not delivered, only fetched. This PR releases the delivery count for records when closing a share session normally. Co-authored-by: d00791190 <dinglan6@huawei.com> Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>, Andrew Schofield <aschofield@confluent.io>	2025-05-01 14:40:03 +01:00
Chirag Wadhwa	800612e4a7	KAFKA-19015: Remove share session from cache on share consumer connection drop (#19329 ) Up till now, the share sessions in the broker were only attempted to evict when the share session cache was full and a new session was trying to get registered. With the changes in this PR, whenever a share consumer gets disconnected from the broker, the corresponding share session would be evicted from the cache. Note - `connectAndReceiveWithoutClosingSocket` has been introduced in `GroupCoordinatorBaseRequestTest`. This method creates a socket connection, sends the request, receives a response but does not close the connection. Instead, these sockets are stored in a ListBuffer `openSockets`, which are closed in tearDown method after each test is run. Also, all the `connectAndReceive` calls in `ShareFetchAcknowledgeRequestTest` have been replaced by `connectAndReceiveWithoutClosingSocket`, because these tests depends upon the persistence of the share sessions on the broker once registered. But, with the new code introduced, as soon as the socket connection is closed, a connection drop is assumed by the broker, leading to session eviction. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>, Andrew Schofield <aschofield@confluent.io>	2025-05-01 14:36:18 +01:00
Lianet Magrans	1059af4eac	MINOR: Improve docs for client group configs (#19605 ) CI / build (push) Waiting to run Details Improve java docs for session and HB interval client configs & fix max.poll.interval description Reviewers: David Jacot <djacot@confluent.io>	2025-04-30 14:04:16 -04:00
Andrew Schofield	ce97b1d5e7	KAFKA-16894: Exploit share feature [3/N] (#19542 ) This PR uses the v1 of the ShareVersion feature to enable share groups for KIP-932. Previously, there were two potential configs which could be used - `group.share.enable=true` and including "share" in `group.coordinator.rebalance.protocols`. After this PR, the first of these is retained, but the second is not. Instead, the preferred switch is the ShareVersion feature. The `group.share.enable` config is temporarily retained for testing and situations in which it is inconvenient to set the feature, but it should really not be necessary, especially when we get to AK 4.2. The aim is to remove this internal config at that point. No tests should be setting `group.share.enable` any more, because they can use the feature (which is enabled in test environments by default because that's how features work). For tests which need to disable share groups, they now set the share feature to v0. The majority of the code changes were related to correct initialisation of the metadata cache in tests now that a feature is used. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-30 13:27:01 +01:00
PoAn Yang	81881dee83	KAFKA-18760: Deprecate Optional<String> and return String from public Endpoint#listener (#19191 ) * Deprecate org.apache.kafka.common.Endpoint#listenerName. * Add org.apache.kafka.common.Endpoint#listener to replace org.apache.kafka.common.Endpoint#listenerName. * Replace org.apache.kafka.network.EndPoint with org.apache.kafka.common.Endpoint. * Deprecate org.apache.kafka.clients.admin.RaftVoterEndpoint#name * Add org.apache.kafka.clients.admin.RaftVoterEndpoint#listener to replace org.apache.kafka.clients.admin.RaftVoterEndpoint#name Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, TaiJuWu <tjwu1217@gmail.com>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, TengYao Chi <frankvicky@apache.org>, Ken Huang <s7133700@gmail.com>, Bagda Parth , Kuan-Po Tseng <brandboat@gmail.com> --------- Signed-off-by: PoAn Yang <payang@apache.org>	2025-04-30 12:15:33 +08:00
Ken Huang	676e0f2ad6	KAFKA-19139 Plugin#wrapInstance should use LinkedHashMap instead of Map (#19519 ) CI / build (push) Waiting to run Details There will be an update to the PluginMetrics#metricName method: the type of the tags parameter will be changed from Map to LinkedHashMap. This change is necessary because the order of metric tags is important 1. If the tag order is inconsistent, identical metrics may be treated as distinct ones by the metrics backend 2. KAFKA-18390 is updating metric naming to use LinkedHashMap. For consistency, we should follow the same approach here. Reviewers: TengYao Chi <frankvicky@apache.org>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, lllilllilllilili	2025-04-30 10:43:01 +08:00
Bill Bejeck	431cffc93f	KAFKA-19135 Migrate initial IQ support for KIP-1071 from feature branch to trunk (#19588 ) This PR is a migration of the initial IQ support for KIP-1071 from the feature branch to trunk. It includes a parameterized integration test that expects the same results whether using either the classic or new streams group protocol. Note that this PR will deliver IQ information in each heartbeat response. A follow-up PR will change that to be only sending IQ information when assignments change. Reviewers Lucas Brutschy <lucasbru@apache.org>	2025-04-29 20:08:49 -04:00
Matthias J. Sax	3bb15c5dee	MINOR: improve JavaDocs for consumer CloseOptions (#19546 ) Reviewers: TengYao Chi <frankvicky@apache.org>, PoAn Yang <payang@apache.org>, Lianet Magrans <lmagrans@confluent.io>, Anna Sophie Blee-Goldman <ableegoldman@apache.org>	2025-04-29 16:38:16 -07:00
Omnia Ibrahim	6f783f8536	KAFKA-10551: Add topic id support to produce request and response (#15968 ) - Add support topicId in `ProduceRequest`/`ProduceResponse`. Topic name and Topic Id will become `ignorable` following the footstep of `FetchRequest`/`FetchResponse` - ReplicaManager still look for `HostedPartition` using `TopicPartition` and doesn't check topic id. This is an [OPEN QUESTION] if we should address this in this pr or wait for [KAFKA-16212](https://issues.apache.org/jira/browse/KAFKA-16212) as this will update `ReplicaManager::getPartition` to use `TopicIdParittion` once we update the cache. Other option is that we compare provided `topicId` with `Partition` topic id and return `UNKNOW_TOPIC_ID` or `UNKNOW_TOPIC_PARTITION` if we can't find partition with matched topic id. Reviewers: Jun Rao <jun@confluent.io>, Justine Olshan <jolshan@confluent.io>	2025-04-29 15:37:28 -07:00
Ritika Reddy	2fdb687029	KAFKA-19082: [2/4] Add preparedTxnState class to Kafka Producer (KIP-939) (#19470 ) CI / build (push) Waiting to run Details This is part of the client side changes required to enable 2PC for KIP-939 New KafkaProducer.PreparedTxnState class is going to be defined as following: ``` static public class PreparedTxnState { public String toString(); public PreparedTxnState(String serializedState); public PreparedTxnState(); } ``` The objects of this class can serialize to / deserialize from a string value and can be written to / read from a database. The implementation is going to store producerId and epoch in the format producerId:epoch Reviewers: Artem Livshits <alivshits@confluent.io>, Justine Olshan <jolshan@confluent.io>	2025-04-29 11:52:02 -07:00
David Jacot	6d67d82d5b	MINOR: Cleanup OffsetFetchRequest/Response in MessageTest (#19576 ) CI / build (push) Waiting to run Details The tests related of OffsetFetch request/response in MessageTest are incomprehensible. This patch rewrites them in a simpler way. Reviewers: TengYao Chi <frankvicky@apache.org>	2025-04-28 13:14:23 +02:00
David Jacot	be194f5dba	MINOR: Simplify OffsetFetchRequest (#19572 ) While working on https://github.com/apache/kafka/pull/19515, I came to the conclusion that the OffsetFetchRequest is quite messy and overall too complicated. This patch rationalize the constructors. OffsetFetchRequest has a single constructor accepting the OffsetFetchRequestData. This will also simplify adding the topic ids. All the changes are mechanical, replacing data structures by others. Reviewers: PoAn Yang <payang@apache.org>, TengYao Chi <frankvicky@apache.org>, Lianet Magran <lmagrans@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-27 18:58:30 +02:00
Chirag Wadhwa	2f9c2dd828	KAFKA-16718-3/n: Added the ShareGroupStatePartitionMetadata record during deletion of share group offsets (#19478 ) This is a follow up PR for implementation of DeleteShareGroupOffsets RPC. This PR adds the ShareGroupStatePartitionMetadata record to __consumer__offsets topic to make sure the topic is removed from the initializedTopics list. This PR also removes partitions from the request and response schemas for DeleteShareGroupState RPC Reviewers: Sushant Mahajan <smahajan@confluent.io>, Andrew Schofield <aschofield@confluent.io>	2025-04-25 22:01:48 +01:00
Ken Huang	b4b80731c1	KAFKA-19042 Move PlaintextConsumerFetchTest to client-integration-tests module (#19520 ) Use Java to rewrite `PlaintextConsumerFetchTest` by new test infra and move it to client-integration-tests module. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-26 00:09:23 +08:00
Lucas Brutschy	732ed0696b	KAFKA-19190: Handle shutdown application correctly (#19544 ) If the streams rebalance protocol is enabled in StreamsUncaughtExceptionHandlerIntegrationTest, the streams application does not shut down correctly upon error. There are two causes for this. First, sometimes, the SHUTDOWN_APPLICATION code only sent with the leave heartbeat, but that is not handled broker side. Second, the SHUTDOWN_APPLICATION code wasn't properly handled client-side at all. Reviewers: Bruno Cadonna <cadonna@apache.org>, Bill Bejeck <bill@confluent.io>, PoAn Yang <payang@apache.org>	2025-04-25 09:56:09 +02:00
PoAn Yang	36d2498fb3	MINOR: Use meaningful name in AsyncKafkaConsumerTest (#19550 ) Replace names like a, b, c, ... with meaningful names in AsyncKafkaConsumerTest. Follow-up: https://github.com/apache/kafka/pull/19457#discussion_r2056254087 Signed-off-by: PoAn Yang <payang@apache.org> Reviewers: Bill Bejeck <bbejeck@apache.org>, Ken Huang <s7133700@gmail.com>	2025-04-24 17:17:33 -04:00
David Jacot	a948537704	MINOR: Small refactor in group coordinator (#19551 ) This patch does a few code changes: * It cleans up the GroupCoordinatorService; * It moves the helper methods to validate request to Utils; * It moves the helper methods to create the assignment for the ConsumerGroupHeartbeatResponse and the ShareGroupHeartbeatResponse from the GroupMetadataManager to the respective classes. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, Jeff Kim <jeff.kim@confluent.io>	2025-04-24 20:57:23 +02:00
Ritika Reddy	62fe528f4b	KAFKA-19082: [1/4] Add client config for enable2PC and overloaded initProducerId (KIP-939) (#19429 ) This is part of the client side changes required to enable 2PC for KIP-939 Producer Config: transaction.two.phase.commit.enable The default would be ‘false’. If set to ‘true’, the broker is informed that the client is participating in two phase commit protocol and transactions that this client starts never expire. Overloaded InitProducerId method If the value is 'true' then the corresponding field is set in the InitProducerIdRequest Reviewers: Justine Olshan <jolshan@confluent.io>, Artem Livshits <alivshits@confluent.io>	2025-04-24 09:41:06 -07:00
Apoorv Mittal	3c05dfdf0e	KAFKA-18889: Make records in ShareFetchResponse non-nullable (#19536 ) This PR marks the records as non-nullable for ShareFetch. This PR is as per the changes for Fetch: https://github.com/apache/kafka/pull/18726 and some work for ShareFetch was done here: https://github.com/apache/kafka/pull/19167. I tested with marking `records` as non-nullable in ShareFetch, which required additional handling. The same has been fixed in current PR. Reviewers: Andrew Schofield <aschofield@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>, TengYao Chi <frankvicky@apache.org>, PoAn Yang <payang@apache.org>	2025-04-24 16:32:08 +01:00
Vikas Singh	f4ab3a2275	MINOR: Use readable interface to parse response (#19353 ) The generated response data classes take Readable as input to parse the Response. However, the associated response objects take ByteBuffer as input and thus convert them to Readable using `new ByteBufferAccessor` call. This PR changes the parse method of all the response classes to take the Readable interface instead so that no such conversion is needed. To support parsing the ApiVersionsResponse twice for different version this change adds the "slice" method to the Readable interface. Reviewers: José Armando García Sancio <jsancio@apache.org>, Truc Nguyen <[trnguyen@confluent.io](mailto:trnguyen@confluent.io)>, Aadithya Chandra <[aadithya.c@gmail.com](mailto:aadithya.c@gmail.com)>	2025-04-24 11:09:06 -04:00
Andrew Schofield	f0f5571dbb	MINOR: Change KIP-932 log messages from early access to preview (#19547 ) Change the log messages which used to warn that KIP-932 was an Early Access feature to say that it is now a Preview feature. This will make the broker logs far less noisy when share groups are enabled. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-24 11:22:17 +01:00
PoAn Yang	3fae785ea0	KAFKA-19110: Add missing unit test for Streams-consumer integration (#19457 ) - Construct `AsyncKafkaConsumer` constructor and verify that the `RequestManagers.supplier()` contains Streams-specific data structures. - Verify that `RequestManagers` constructs the Streams request managers correctly - Test `StreamsGroupHeartbeatManager#resetPollTimer()` - Test `StreamsOnTasksRevokedCallbackCompletedEvent`, `StreamsOnTasksAssignedCallbackCompletedEvent`, and `StreamsOnAllTasksLostCallbackCompletedEvent` in `ApplicationEventProcessor` - Test `DefaultStreamsRebalanceListener` - Test `StreamThread`. - Test `handleStreamsRebalanceData`. - Test `StreamsRebalanceData`. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>, Bill Bejeck <bill@confluent.io> Signed-off-by: PoAn Yang <payang@apache.org>	2025-04-24 10:38:22 +02:00
Kirk True	8b4560e3f0	KAFKA-15767 Refactor TransactionManager to avoid use of ThreadLocal (#19440 ) Introduces a concrete subclass of `KafkaThread` named `SenderThread`. The poisoning of the TransactionManager on invalid state changes is determined by looking at the type of the current thread. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-04-24 00:31:30 +08:00
Bruno Cadonna	efd785274e	KAFKA-19124: Follow up on code improvements (#19453 ) Improves a variable name and handling of an Optional. Reviewers: Bill Bejeck <bill@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>, PoAn Yang <payang@apache.org>	2025-04-23 14:24:33 +02:00
David Jacot	71d08780d1	KAFKA-14690; Add TopicId to OffsetCommit API (#19461 ) This patch extends the OffsetCommit API to support topic ids. From version 10 of the API, topic ids must be used. Originally, we wanted to support both using topic ids and topic names from version 10 but it turns out that it makes everything more complicated. Hence we propose to only support topic ids from version 10. Clients which only support using topic names can either lookup the topic ids using the Metadata API or stay on using an earlier version. The patch only contains the server side changes and it keeps the version 10 as unstable for now. We will mark the version as stable when the client side changes are merged in. Reviewers: Lianet Magrans <lmagrans@confluent.io>, PoAn Yang <payang@apache.org>	2025-04-23 08:22:09 +02:00
Andrew Schofield	e78e106221	MINOR: Improve javadoc for share consumer (#19533 ) Small improvements to share consumer javadoc. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-22 15:54:05 +01:00
Andrew Schofield	66147d5de7	KAFKA-19057: Stabilize KIP-932 RPCs for AK 4.1 (#19378 ) This PR removes the unstable API flag for the KIP-932 RPCs. The 4 RPCs which were exposed for the early access release in AK 4.0 are stabilised at v1. This is because the RPCs have evolved over time and AK 4.0 clients are not compatible with AK 4.1 brokers. By stabilising at v1, the API version checks prevent incompatible communication and server-side exceptions when trying to parse the requests from the older clients. Reviewers: Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-22 11:43:32 +01:00
Rich Chen	ae771d73d1	KAFKA-8830 make Record Headers available in onAcknowledgement (#17099 ) Two sets of tests are added: 1. KafkaProducerTest - when send success, both record.headers() and onAcknowledgement headers are read only - when send failure, record.headers() is writable as before and onAcknowledgement headers is read only 2. ProducerInterceptorsTest - make both old and new onAcknowledgement method are called successfully Reviewers: Lianet Magrans <lmagrans@confluent.io>, Omnia Ibrahim <o.g.h.ibrahim@gmail.com>, Matthias J. Sax <matthias@confluent.io>, Andrew Schofield <aschofield@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-21 21:01:55 +08:00
Hong-Yi Chen	8fa0d9723f	MINOR: Fix typo in ApiKeyVersionsProvider exception message (#19521 ) This patch addresses issue #19516 and corrects a typo in `ApiKeyVersionsProvider`: when `toVersion` exceeds `latestVersion`, the `IllegalArgumentException` message was erroneously formatted with `fromVersion`. The format argument has been updated to use `toVersion` so that the error message reports the correct value. Reviewers: Ken Huang <s7133700@gmail.com>, PoAn Yang <payang@apache.org>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-21 15:35:47 +08:00
David Jacot	b94c7f9167	MINOR: Extend @ApiKeyVersionsSource annotation (#19516 ) This patch extends the `@ApiKeyVersionsSource` annotation to allow specifying the `fromVersion` and the `toVersion`. This is pretty handy when we only want to test a subset of the versions. Reviewers: Kuan-Po Tseng <brandboat@gmail.com>, TengYao Chi <kitingiao@gmail.com>	2025-04-20 12:25:27 +08:00
Matthias J. Sax	810beef50e	MINOR: improve (De)Serializer JavaDocs (#19467 ) Reviewers: Kirk True <ktrue@confluent.io>, Lianet Magrans <lmagrans@confluent.io>	2025-04-17 11:23:15 -07:00
Logan Zhu	c6496e0c57	MINOR: Cleanup 0.10.x legacy references in ClusterResourceListener and TopicConfig (clients module) (#19388 ) This PR is a minor follow-up to [PR #19320](https://github.com/apache/kafka/pull/19320), which cleaned up 0.10.x legacy information from the clients module. It addresses remaining reviewer suggestions that were not included in the original PR: - `ClusterResourceListener`: Removed "Note the minimum supported broker version is 2.1." per review suggestion to avoid repeating version-specific details across multiple classes. - `TopicConfig`: Simplified `MAX_MESSAGE_BYTES_DOC` by removing obsolete notes about behavior in versions prior to 0.10.2. These changes help reduce outdated version information in client documentation and improve clarity. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-17 23:17:42 +08:00
Andrew Schofield	8d66481a83	KAFKA-17897 Deprecate Admin.listConsumerGroups (#19477 ) The final part of KIP-1043 is to deprecate Admin.listConsumerGroups() in favour of Admin.listGroups() which works for all group types. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-17 23:00:57 +08:00
Lucas Brutschy	5f80de3923	KAFKA-19162: Topology metadata contains non-deterministically ordered topic configs (#19491 ) Topology description sent to broker in KIP-1071 contains non-deterministically ordered topic configs. Since the topology is compared to the groups topology upon joining we may run into `INVALID_REQUEST: Topology updates are not supported yet` failures if the topology sent by the application does not match the group topology due to different topic config order. This PR ensures that topic configs are ordered, to avoid an `INVALID_REQUEST` error. Reviewers: Matthias J. Sax <matthias@confluent.io>	2025-04-16 21:12:17 -07:00
yunchi	effbad9e80	KAFKA-19151 docs: clarify that flush.ms requires log.flush.scheduler.interval.ms config (#19479 ) Enhanced docs of `flush.ms` to remind users the flush is triggered by `log.flush.scheduler.interval.ms`. Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-17 11:19:44 +08:00
TaiJuWu	23e7158665	KAFKA-19002 Rewrite ListOffsetsIntegrationTest and move it to clients-integration-test (#19460 ) the following tasks should be addressed in this ticket rewrite it by 1. new test infra 2. use java 3. move it to clients-integration-test Reviewers: TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-17 02:26:23 +08:00
Andrew Schofield	6a4207f12a	KAFKA-19158: Add SHARE_SESSION_LIMIT_REACHED error code (#19492 ) Add the new `SHARE_SESSION_LIMIT_REACHED` error code which is used when an attempt is made to open a new share session when the share session limit of the broker has already been reached. Support in the client and broker will follow in subsequent PRs. Reviewers: Lianet Magrans <lmagrans@confluent.io>	2025-04-16 18:00:07 +01:00
David Jacot	6e26ec06bb	MINOR: Update GroupCoordinator interface to use AuthorizableRequestContext instead of RequestContext (#19485 ) This patch updates the `GroupCoordinator` interface to use `AuthorizableRequestContext` instead of using `RequestContext`. It makes the interface more generic. The only downside is that the request version in `AuthorizableRequestContext` is an `int` instead of a `short` so we had to adapt it in a few places. We opted for using `int` directly wherever possible. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, Rajini Sivaram <rajinisivaram@googlemail.com>	2025-04-16 09:12:11 -07:00
Ken Huang	ae608c1cb2	KAFKA-19042 Move PlaintextConsumerCallbackTest to client-integration-tests module (#19298 ) Use Java to rewrite `PlaintextConsumerCallbackTest` by new test infra and move it to client-integration-tests module. Reviewers: TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-16 11:57:14 +08:00
Mickael Maison	fb2ce76b49	KAFKA-18888: Add KIP-877 support to Authorizer (#19050 ) This also adds metrics to StandardAuthorizer Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, Ken Huang <s7133700@gmail.com>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>, TaiJuWu <tjwu1217@gmail.com>	2025-04-15 19:40:24 +02:00
Ritika Reddy	598eb13d07	KAFKA-15370: ACL changes to support 2PC (KIP-939) (#19364 ) This patch adds ACL support for 2PC as a part of KIP-939 A new value will be added to the enum AclOperation: TWO_PHASE_COMMIT ((byte) 15 . When InitProducerId comes with enable2Pc=true, it would have to have both WRITE and TWO_PHASE_COMMIT operation enabled on the transactional id resource. The kafka-acls.sh tool is going to support a new --operation TwoPhaseCommit. Reviewers: Artem Livshits <alivshits@confluent.io>, PoAn Yang <poan.yang@suse.com>, Justine Olshan <jolshan@confluent.io>	2025-04-15 08:39:46 -07:00
Shivsundar R	f737ef31d9	KAFKA-18900: Implement share.acknowledgement.mode to choose acknowledgement mode (#19417 ) Choose the acknowledgement mode based on the config (`share.acknowledgement.mode`) and not on the basis of how the user designs the application. - The default value of the config is `IMPLICIT`, so if any empty/null/invalid value is configured, then the mode defaults to `IMPLICIT`. - Removed AcknowledgementModes `UNKNOWN` and `PENDING` as they are no longer required. - Added code to ensure if the application has any unacknowledged records in a batch in "`explicit`" mode, then it will throw an `IllegalStateException`. The expectation is if the mode is "explicit", all the records received in that `poll()` would be acknowledged before the next call to `poll()`. - Modified the `ConsoleShareConsumer` to configure the mode to "explicit" as it was using the explicit mode of acknowledging records. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-04-15 16:38:33 +01:00
Shivsundar R	6c3995b954	MINOR: Port changes from KAFKA-18569 for ShareConsumers (#19402 ) ShareConsumers` may wait on an unneeded `FindCoordinator` during `close()`(i.e after the acknowledgements are sent). https://github.com/apache/kafka/pull/18590 added the `StopFindCoordinatorOnClose` event and was used by the regular consumers. We are using the same event in `ShareConsumers` as well to prevent sending this event when coordinator is no longer needed. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-04-15 16:04:22 +01:00
Xuan-Zhang Gong	c527530e80	KAFKA-19042 Move ProducerCompressionTest, ProducerFailureHandlingTest, and ProducerIdExpirationTest to client-integration-tests module (#19319 ) include three test case - ProducerCompressionTest - ProducerFailureHandlingTest - ProducerIdExpirationTest Reviewers: Ken Huang <s7133700@gmail.com>, PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-15 16:34:47 +08:00
Azhar Ahmed	4cdd4b617c	KAFKA-19071: Fix doc for remote.storage.enable (#19345 ) As of 3.9, Kafka allows disabling remote storage on a topic after it was enabled. It allows subsequent enabling and disabling too. However the documentation says otherwise and needs to be corrected. Doc: https://kafka.apache.org/39/documentation/#topicconfigs_remote.storage.enable Reviewers: Luke Chen <showuon@gmail.com>, PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>	2025-04-14 11:08:49 +08:00
PoAn Yang	34a87d3477	KAFKA-19042 Move TransactionsWithMaxInFlightOneTest to client-integration-tests module (#19289 ) Use Java to rewrite `TransactionsWithMaxInFlightOneTest` by new test infra and move it to client-integration-tests module. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-04-11 12:04:19 +08:00
Jhen-Yung Hsu	90e7b53799	MINOR: Remove unused `ApiVersions` variable from Sender and RecordAccumulator (#19399 ) Remove unused `ApiVersions` variable from Sender and RecordAccumulator. Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, Parker Chang <parkerhiphop027@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-11 11:23:41 +08:00
Kaushik Raina	b3ba7bc929	KAFKA-18782: Extend ApplicationRecoverableException related exceptions (#19354 ) Summary Extend ApplicationRecoverableException related exceptions Reviewers: Artem Livshits <alivshits@confluent.io>, Justine Olshan <jolshan@confluent.io>	2025-04-10 16:57:28 -07:00
Bruno Cadonna	c11938c926	KAFKA-19124: Use consumer background event queue for Streams events (#19421 ) In the first version of the integration of the stream thread with the new Streams rebalance protocol, the consumer used a dedicated event queue for Streams/specific background events to request the stream thread to call the rebalance callbacks. That led to an issue where the consumer times out when unsubscribing. This commit gets rid of the dedicated queue and incorporates the Streams-specific background events into event queue used by the consumer. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-04-10 21:06:06 +02:00
TengYao Chi	b649b1ed5d	KAFKA-18935: Ensure brokers do not return null records in FetchResponse (#19167 ) JIRA: KAFKA-18935 This patch ensures the broker will not return null records in FetchResponse. For more details, please refer to the ticket. Reviewers: Ismael Juma <ismael@juma.me.uk>, Chia-Ping Tsai <chia7712@gmail.com>, Jun Rao <junrao@gmail.com>	2025-04-10 22:21:00 +08:00
Abhinav Dixit	699ae1b75b	KAFKA-16729: Support isolation level for share consumer (#19261 ) This PR adds the share group dynamic config `share.isolation.level`. Until now, share groups only supported `READ_UNCOMMITTED` isolation level type. With this PR, we aim to support `READ_COMMITTED` isolation type to share groups. Reviewers: Andrew Schofield <aschofield@confluent.io>, Jun Rao <junrao@gmail.com>, Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-10 09:00:03 +01:00
Florian Hussonnois	eeb1214ba8	KAFKA-18962: Fix onBatchRestored call in GlobalStateManagerImpl (#19188 ) Call the StateRestoreListener#onBatchRestored with numRestored and not the totalRestored when reprocessing state See: https://issues.apache.org/jira/browse/KAFKA-18962 Reviewers: Anna Sophie Blee-Goldman <ableegoldman@apache.org>, Matthias Sax <mjsax@apache.org>	2025-04-09 13:17:38 -07:00
Bruno Cadonna	2a370ed721	KAFKA-19037: Integrate consumer-side code with Streams (#19377 ) The consumer adaptations for the new Streams rebalance protocol need to be integrated into the Streams code. This commit does the following: - creates an async Kafka consumer - with a Streams heartbeat request manager - with a Streams membership manager - integrates consumer code with the Streams membership manager and the Streams heartbeat request manager - processes the events from the consumer network thread (a.k.a. background thread) that request the invocation of the "on tasks revoked", "on tasks assigned", and "on all tasks lost" callbacks - executes the callbacks - sends to the consumer network thread the events signalling the execution of the callbacks - adapts SmokeTestDriverIntegrationTest to use the new Streams rebalance protocol This commit misses some unit test coverage, but it also unblocks other work on trunk regarding the new Streams rebalance protocol. The missing unit tests will be added soon. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-04-09 13:26:51 +02:00
Chirag Wadhwa	5148174196	KAFKA-16718-2/n: KafkaAdminClient and GroupCoordinator implementation for DeleteShareGroupOffsets RPC (#18976 ) This PR contains the implementation of KafkaAdminClient and GroupCoordinator for DeleteShareGroupOffsets RPC. - Added `deleteShareGroupOffsets` to `KafkaAdminClient` - Added implementation for `handleDeleteShareGroupOffsetsRequest` in `KafkaApis.scala` - Added `deleteShareGroupOffsets` to `GroupCoordinator` as well. internally this makes use of `persister.deleteState` to persist the changes in share coordinator Reviewers: Andrew Schofield <aschofield@confluent.io>, Sushant Mahajan <smahajan@confluent.io>	2025-04-09 07:31:06 +01:00
lorcan	434b0d39ae	MINOR: use enum map for error counts map (#19314 ) Java provides a specialised Map where Enums are the keys, which can provide some performance improvements. https://docs.oracle.com/javase/8/docs/api/java/util/EnumMap.html I have updated the Java code where possible to use an EnumMap rather than a HashMap and run the unit tests under the requests directory. Reviewers: Matthias J. Sax <matthias@confluent.io>, Lianet Magrans <lmagrans@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-09 02:01:02 +08:00
Ken Huang	2f086d188f	KAFKA-18892: Add KIP-877 support for ClientQuotaCallback (#19068 ) Allow ClientQuotaCallback to implement Monitorable and register metrics. Reviewers: Mickael Maison <mickael.maison@gmail.com>, TaiJuWu <tjwu1217@gmail.com>, Jhen-Yung Hsu <jhenyunghsu@gmail.com>	2025-04-08 16:58:29 +02:00
Nick Guo	fcf6da0a0d	KAFKA-19098 Remove `lastOffset` from PartitionResponse (#19398 ) The `lastOffset` is not used actually, so it can be removed. Reviewers: Jhen-Yung Hsu <jhenyunghsu@gmail.com>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-08 00:06:02 +08:00
Shivsundar R	2d02f1d52d	KAFKA-19084: Port KAFKA-16224, KAFKA-16764 for ShareConsumers (#19369 ) Currently for ShareConsumers, if we receive an `UNKNOWN_TOPIC_OR_PARTITION` error code in the `ShareAcknowledgeResponse`, then we retry sending the acknowledgements until the timer expires. We ideally do not want this when a topic/partition is deleted, hence like the `CommitRequestManager`(https://github.com/apache/kafka/pull/15581), we will treat this error as fatal and not retry the acknowledgements. PR also suppresses `InvalidTopicException` during `unsubscribe()` which was also added in the `AsyncKafkaConsumer`(https://github.com/apache/kafka/pull/16043). It was later removed in the regular consumer as they notified the background operations of metadata errors instead of propagating them via `ErrorEvent`. `ShareConsumerImpl` however does not require that change and it still propagates the metadata error back to the application. So we would need to suppress this exception during unsubscribe(). Reviewers: Andrew Schofield <aschofield@confluent.io>, Sushant Mahajan <smahajan@confluent.io>	2025-04-07 10:04:48 +01:00
Hong-Yi Chen	6dd2cc70c3	MINOR: Clean up comments and remove unused code in RecordVersion and CreateTopicsRequestTest (#19342 ) ## Summary This PR updates the `RecordVersion` javadoc for clarity. It removes outdated references to `message.format.version` mentioned in the [Kafka 4.0 upgrade documentation](`48f06981ee/40/upgrade.html (L135)`) and aligns with feedback from a previous discussion in [#19325 ](https://github.com/apache/kafka/pull/19325). ## Changes - Cleaned up javadoc in `RecordVersion` - Removed outdated or deprecated references Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-07 07:47:06 +08:00
Thomas Gebert	a65626b6a8	MINOR: Add functionalinterface to the producer callback (#19366 ) The Callback interface is a perfect example of a place that can use the functionalinterface in Java. Strictly for Java, this isn't "required" since Java will automatically coerce, but for Clojure (and other JVM languages I belive) to interop with Java lambdas it needs the FunctionalInterface annotation. Since FunctionalInterface doesn't add any overhead and provides compiler-enforced documentation, I don't see any reason not to have this. This has already been added into Kafka Streams here: https://github.com/apache/kafka/pull/19234#pullrequestreview-2740742487 I am happy to add it to any other spots in that might be useful too. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-04-06 22:21:09 +08:00
Parker Chang	9f676dd7e2	MINOR: Clean up unreachable code in FetcherTest (#19376 ) This is from [#16532's comment](https://github.com/apache/kafka/pull/16532/files#r2028985028): The forEach loop in the assertion will never execute because `nonResponseData` is empty. This happens because the above assertion `emptyMap()` has a size of 0, so there are no elements to iterate over. Reviewers: PoAn Yang <payang@apache.org>, Ken Huang <s7133700@gmail.com>, TaiJuWu <tjwu1217@gmail.com>, TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-06 22:17:02 +08:00
TengYao Chi	74acbd200d	KAFKA-16758: Extend Consumer#close with an option to leave the group or not (#17614 ) JIRA: [KAFKA-16758](https://issues.apache.org/jira/browse/KAFKA-16758) This PR is aim to deliver [KIP-1092](https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=321719077), please refer to KIP-1092 and KAFKA-16758 for further details. Reviewers: Anna Sophie Blee-Goldman <ableegoldman@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>, Kirk True <kirk@kirktrue.pro>	2025-04-05 22:02:45 -07:00
PoAn Yang	3d96b20630	KAFKA-19042 Move TransactionsExpirationTest to client-integration-tests module (#19288 ) Use Java to rewrite `TransactionsExpirationTest` by new test infra and move it to client-integration-tests module. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-04-05 20:01:31 +08:00
TaiJuWu	ebb62812d9	KAFKA-19074 Remove the cached responseData from ShareFetchResponse (#19357 ) Jira: https://issues.apache.org/jira/browse/KAFKA-19074 Similar fix https://github.com/apache/kafka/pull/16532 `2b8aff58b5` make it accept input to return "partial" data. The content of output is based on the input but we cache the output ... It will return same output even though we pass different input. That is a potential bug. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-05 19:56:59 +08:00
Andrew Schofield	d4d9f11816	KAFKA-18761: [2/N] List share group offsets with state and auth (#19328 ) This PR approaches completion of Admin.listShareGroupOffsets() and kafka-share-groups.sh --describe --offsets. Prior to this patch, kafka-share-groups.sh was only able to describe the offsets for partitions which were assigned to active members. Now, the Admin.listShareGroupOffsets() uses the persister's knowledge of the share-partitions which have initialised state. Then, it uses this list to obtain a complete set of offset information. The PR also implements the topic-based authorisation checking. If Admin.listShareGroupOffsets() is called with a list of topic-partitions specified, the authz checking is performed on the supplied list, returning errors for any topics to which the client is not authorised. If Admin.listShareGroupOffsets() is called without a list of topic-partitions specified, the list of topics is discovered from the persister as described above, and then the response is filtered down to only show the topics to which the client is authorised. This is consistent with other similar RPCs in the Kafka protocol, such as OffsetFetch. Reviewers: David Arthur <mumrah@gmail.com>, Sushant Mahajan <smahajan@confluent.io>, Apoorv Mittal <apoorvmittal10@gmail.com>	2025-04-04 13:25:19 +01:00
Logan Zhu	a4375045d6	KAFKA-19055 Cleanup the 0.10.x information from clients module (#19320 ) Removes outdated references to Kafka 0.10.x in the clients module documentation. Since the baseline version is now 2.1, any mentions of versions earlier than this are unnecessary and have been removed or updated accordingly. Changes: - Updated `ClusterResource`, `ClusterResourceListener`, and `DescribeClusterResult` Javadoc to reflect the minimum supported broker version as 2.1. - Updated `TopicConfig` documentation: Removed references to consumers older than 0.10.2. - Removed references to 0.10.x and adjusted explanations to remain relevant for newer versions. Testing & Impact: - This PR only modifies Javadoc comments—no functional code changes. - No impact on existing functionality. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-04-04 04:17:13 +08:00
Thomas Gebert	db4e74b46e	MINOR: Add Functional Interface annotation to interfaces used by Lambdas (#19234 ) Adds the FunctionalInterface annotation to relevant Kafka Streams classes. While this is not strictly required for Java, it's still best practice and also useful for better integration with other JVM languages, for example Clojure, to allow using these interfaces as lambdas. Reviewers: Matthias J. Sax <matthias@confluent.io>	2025-04-03 09:30:56 -07:00
Ritika Reddy	eeffd8c475	KAFKA-19003: Add forceTerminateTransaction command to CLI tools (#19276 ) This patch is part of KIP-939 [Support Participation in 2PC](https://cwiki.apache.org/confluence/display/KAFKA/KIP-939%3A+Support+Participation+in+2PC) The kafka-transactions.sh tool will support a new command --forceTerminateTransaction It has one required argument --transactionalId that would take the transactional id for the transaction to be terminated. The command uses the existing Admin#fenceProducers method to forcefully abort the transaction associated with the specified transactional ID. Under the hood, it sends an InitProducerId request to the transaction coordinator with the given transactional ID and keepPreparedTxn = false by default. This is aligned with the functionality outlined in the KIP. We will be creating a new public method in the Admin Client public TerminateTransactionResult forceTerminateTransaction(String transactionalId), and re-use the existing fence producer method. Reviewers: Artem Livshits <alivshits@confluent.io>, Justine Olshan <jolshan@confluent.io>	2025-04-02 11:51:26 -07:00
Andrew Schofield	cee55dbdec	KAFKA-18794: Disable flaky tests pending investigation (#19340 ) KafkaShareConsumerTest is proving very flaky. The behaviour of MockClient does not appear to match the expectations of the test. This PR disables the flaky tests to reduce build noise. When a proper solution has been worked out, the tests can be re-enabled. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-04-01 16:52:06 +01:00
Shivsundar R	e301508b53	MINOR: Add check in ShareConsumerImpl to send acknowledgements of control records when ShareFetch is empty. (#19295 ) Currently if we received just a control record in the `ShareFetchResponse`, then the currentFetch in `ShareConsumerImpl` would not be updated as the record is ignored. But in the process, we lose the acknowledgment for this control record which is a GAP. PR fixes this by adding an additional map for control record acknowledgements in `ShareFetchEvent`. This updates both the ShareConsumerImpl and ShareConsumeRequestManager to accommodate the additional map. Added a unit test in `ShareConsumerImplTest` and `ShareConsumeRequestManagerTest` to verify the changes. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-04-01 14:15:03 +01:00
Shivsundar R	ed77397814	KAFKA-19062: Port changes from KAFKA-18645 to share-consumers (#19335 ) Limits waiting when closing a share consumer to request.timeout.ms. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-04-01 13:08:19 +01:00
Apoorv Mittal	4aa81204ff	KAFKA-19018,KAFKA-19063: Implement maxRecords and acquisition lock timeout in share fetch request and response resp. (#19334 ) PR add `MaxRecords` to share fetch request and also adds `AcquisitionLockTimeout` to share fetch response. PR also removes internal broker config of `max.fetch.records`. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-04-01 12:23:06 +01:00
Ismael Juma	b375bb099b	MINOR: Remove unused `ApiKeys.minRequiredInterBrokerMagic` (#19325 ) Reviewers: David Jacot <david.jacot@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-31 10:41:05 -07:00
TengYao Chi	20546930ae	KAFKA-19042 Move ConsumerTopicCreationTest to client-integration-tests module (#19283 ) This patch moves `ConsumerTopicCreationTest` to the `client-integration-tests` and rewrite it as Java. The patch also streamlines the test flow. In the Scala version, there is a producer that produces messages, but this is not the main purpose of the `ConsumerTopicCreationTest`. Reviewers: Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-31 20:15:54 +08:00
Kuan-Po Tseng	c095faa578	KAFKA-18945 Enhance the docs for Admin APIs (#19315 ) Enhance the documentation for Admin#describeCluster and Admin#describeConfigs to clarify their behavior when using bootstrap.controllers and bootstrap.servers. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-31 13:49:04 +08:00
Nick Guo	c771116b89	KAFKA-19005 improve the documentation of DescribeTopicsOptions#partitionSizeLimitPerResponse (#19268 ) jira: https://issues.apache.org/jira/browse/KAFKA-19005 This PR includes following changes: 1. refine the format 2. highligh that it is supported by topic names <img width="857" alt="999" src="https://github.com/user-attachments/assets/6eec9e2f-b839-430c-b111-2be3a8538593" /> Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-29 03:21:10 +08:00
Nick Guo	9292a22606	KAFKA-19049 Remove the `@ExtendWith(ClusterTestExtensions.class)` from code base (#19299 ) jira: https://issues.apache.org/jira/browse/KAFKA-19049 [KAFKA-18617](https://issues.apache.org/jira/browse/KAFKA-18617) introduced the mechanism to inject the cluster test at runtime, so the integration tests don't need to use `@ExtendWith(ClusterTestExtensions.class)` any more. Reviewers: PoAn Yang <payang@apache.org>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-29 02:15:16 +08:00
Lucas Brutschy	2267902b40	MINOR: Mark streams RPCs as unstable (#19292 ) Streams groups RPCs are not enabled by default, but they should also be marked as unstable. Reviewers: Bruno Cadonna <cadonna@apache.org>	2025-03-27 14:22:01 +01:00
Sushant Mahajan	eb88e78373	KAFKA-18827: Initialize share group state group coordinator impl. [3/N] (#19026 ) * This PR adds impl for the initialize share groups call from the Group Coordinator perspective. * The initialize call on persister instance will be invoked by the `GroupCoordinatorService`, based on the response of the `GroupCoordinatorShard.shareGroupHeartbeat`. If there is new topic subscription or member assignment change (topic paritions incremented), the delta share partitions corresponding to the share group in question are returned as an optional initialize request. * The request is then sent to the share coordinator as an encapsulated timer task because we want the heartbeat response to go asynchronously. * Tests have been added for `GroupCoordinatorService` and `GroupMetadataManager`. Existing tests have also been updated. * A new formatter `ShareGroupStatePartitionMetadataFormatter` has been added for debugging. Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-03-26 19:40:23 +00:00
Vikas Singh	56d1dc1b6e	MINOR: Use readable interface to parse requests (#19163 ) The generated request data type's constructors take Readable as an input. However, the parse method in the AbstractRequest takes a ByteBuffer as input. So to create the corresponding request data objects, each individual concrete Request classes wraps the ByteBuffer into a ByteBufferAccessor. This is boilerplate code present in all the concrete request classes. This changes AbstractRequest's parse method so that subclasses can simply pass the `Readable` they get directly to request data classes. The same change is made to the serialize method to maintain symmetry. Reviewers: Ismael Juma <ismael@juma.me.uk>, José Armando García Sancio <jsancio@apache.org>, Artem Livshits <alivshits@confluent.io>, Truc Nguyen <trnguyen@confluent.io>	2025-03-26 10:13:13 -04:00
Shivsundar R	91758cc99d	KAFKA-18899: Improve handling of timeouts for commitAsync() in ShareConsumer. (#19192 ) Previously, the `ShareConsumer.commitAsync()` method retried sending `ShareAcknowledge` requests indefinitely. Now it will instead use the defaultApiTimeout config to expire the request so that it does not retry forever. PR also fixes a bug in processing `commitSync() `requests, where we need an additional check if the node is free. Co-authored-by: Andrew Schofield <aschofield@confluent.io> Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-03-26 09:06:59 +00:00
ClarkChen	1547204baa	KAFKA-18914 Migrate ConsumerRebootstrapTest to use new test infra (#19154 ) Migrate ConsumerRebootstrapTest to the new test infra and remove the old Scala test. The PR changed three things. * Migrated `ConsumerRebootstrapTest` to new test infra and removed the old Scala test. * Updated the original test case to cover rebootstrap scenarios. * Integrated `ConsumerRebootstrapTest` into `ClientRebootstrapTest` in the `client-integration-tests` module. * Removed the `RebootstrapTest.scala`. Default `ConsumerRebootstrap` config: > properties.put(CommonClientConfigs.METADATA_RECOVERY_STRATEGY_CONFIG, "rebootstrap"); properties.put(CommonClientConfigs.METADATA_RECOVERY_REBOOTSTRAP_TRIGGER_MS_CONFIG, "300000"); properties.put(CommonClientConfigs.SOCKET_CONNECTION_SETUP_TIMEOUT_MS_CONFIG, "10000"); properties.put(CommonClientConfigs.SOCKET_CONNECTION_SETUP_TIMEOUT_MAX_MS_CONFIG, "30000"); properties.put(CommonClientConfigs.RECONNECT_BACKOFF_MS_CONFIG, "50L"); properties.put(CommonClientConfigs.RECONNECT_BACKOFF_MAX_MS_CONFIG, "1000L"); The test case for the consumer with enabled rebootstrap ![Screenshot 2025-03-22 at 9 48 13 PM](https://github.com/user-attachments/assets/8470549f-a24c-43fa-ae44-789cbf422a63) The test case for the consumer with disabled rebootstrap ![Screenshot 2025-03-22 at 9 47 22 PM](https://github.com/user-attachments/assets/0a183464-6a74-449f-8e71-d641a6ea5bb1) Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-26 01:53:42 +08:00
Bruno Cadonna	96196bb03b	KAFKA-18736: Add pollOnClose() and maximumTimeToWait() (#19233 ) Adds pollOnClose() and maximumTimeToWait() to the Streams group heartbeat request manager. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-03-25 09:09:13 +01:00
Bruno Cadonna	266532f562	KAFKA-18736: Handle errors in the Streams group heartbeat request manager (#19230 ) This commit adds error handling to the Streams heartbeat request manager. Errors can occur while sending a heartbeat request and when a response with an error code that is not NONE is received. Some errors are handled explicitly to recover from them or to log specific messages. All the others are handled as fatal errors. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-03-24 21:26:14 +01:00
TaiJuWu	a524fc64b4	MINOR: leverage preProcessParsedConfig within AbstractConfig (#19259 ) In past, we have `AbstractConfig#preProcessParsedConfig` but did not use its return value Reviewers: Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-24 01:19:20 +08:00
ClarkChen	fef9aebb19	KAFKA-18276 Migrate ProducerRebootstrapTest to new test infra (#19046 ) The PR changed three things. * Migrated `ProducerRebootstrapTest` to new test infra and removed the old Scala test. * Updated the original test case to cover rebootstrap scenarios. * Integrated `ProducerRebootstrapTest` into `ClientRebootstrapTest` in the `client-integration-tests` module. Default `ProducerRebootstrap` config: > properties.put(CommonClientConfigs.METADATA_RECOVERY_STRATEGY_CONFIG, "rebootstrap"); properties.put(CommonClientConfigs.METADATA_RECOVERY_REBOOTSTRAP_TRIGGER_MS_CONFIG, "300000"); properties.put(CommonClientConfigs.SOCKET_CONNECTION_SETUP_TIMEOUT_MS_CONFIG, "10000"); properties.put(CommonClientConfigs.SOCKET_CONNECTION_SETUP_TIMEOUT_MAX_MS_CONFIG, "30000"); properties.put(CommonClientConfigs.RECONNECT_BACKOFF_MS_CONFIG, "50L"); properties.put(CommonClientConfigs.RECONNECT_BACKOFF_MAX_MS_CONFIG, "1000L"); The test case for the producer with enabled rebootstrap <img width="1549" alt="Screenshot 2025-03-17 at 10 46 03 PM" src="https://github.com/user-attachments/assets/547840a6-d79d-4db4-98c0-9b05ed04cf60" /> The test case for the producer with disabled rebootstrap <img width="1552" alt="Screenshot 2025-03-17 at 10 46 47 PM" src="https://github.com/user-attachments/assets/2248e809-d9d5-4f3b-a24f-ba1aa0fef728" /> Reviewers: TengYao Chi <kitingiao@gmail.com>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-24 01:09:17 +08:00
Ken Huang	68ecb7720f	MINOR: add log4j2.yaml to clients-integration-tests module (#19252 ) `clients-integration-tests` modules doesn't have the `log4j2.yaml` to setting log, thus we should add. Reviewers: TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-22 02:19:54 +08:00
TaiJuWu	79fe1305b6	KAFKA-18893: Add KIP-877 support to ReplicaSelector (#19064 ) ReplicaSelector implementations can implement Monitorable to register their own metrics. Reviewers: Mickael Maison <mickael.maison@gmail.com>, Ken Huang <s7133700@gmail.com>	2025-03-21 15:39:50 +01:00
David Arthur	8fa3856473	MINOR Mar 19 flaky tests (#19248 ) CoordinatorRequestManagerTest#testMarkCoordinatorUnknownLoggingAccuracy has become flaky again. Last 30 days report shows a sudden re-occurrence https://develocity.apache.org/scans/tests?search.relativeStartTime=P28D&search.rootProjectNames=kafka&search.tags=github,trunk,not:flaky,not:new&search.tasks=test&search.timeZoneId=America%2FNew_York&tests.container=org.apache.kafka.clients.consumer.internals.CoordinatorRequestManagerTest&tests.sortField=FLAKY# Also mark QuorumControllerTest.testMinIsrUpdateWithElr as flaky. Reviewers: Lianet Magrans <lmagrans@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-21 09:26:13 -04:00
David Jacot	0c5e5c5d2d	KAFKA-18329; [2/3] Delete old group coordinator (KIP-848) (#19251 ) This patch is the second of a series of patches to remove the old group coordinator. With the release of Apache Kafka 4.0, the so-called new group coordinator is the default and only option available now. The patch removes `group.coordinator.new.enable` (internal config) and all its usages (integration tests, unit tests, etc.). It also cleans up `KafkaApis` to remove logic only used by the old group coordinator. Reviewers: Jeff Kim <jeff.kim@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-21 05:54:41 -07:00
Ken Huang	e21c46a504	MINOR: Move FileRecord JavaDoc to comment (#19257 ) See: https://github.com/apache/kafka/pull/19214#discussion_r2005945059 Move explaination from Javadoc to comment. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-21 13:47:56 +08:00
Ken Huang	31e1a57c41	KAFKA-18989 Optimize FileRecord#searchForOffsetWithSize (#19214 ) The `lastOffset` includes the entire batch header, so we should check `baseOffset` instead. To optimize this, we need to update the search logic. The previous approach simply checked whether each batch's `lastOffset()` was greater than or equal to the target offset. Once it found the first batch that met this condition, it returned that batch immediately. Now that we are using `baseOffset()`, we need to handle a special case: if the `targetOffset` falls between the `lastOffset` of the previous batch and the `baseOffset` of the matching batch, we should select the matching batch. The updated logic is structured as follows: 1. First, if baseOffset exactly equals targetOffset, return immediately. 2. If we find the first batch with baseOffset greater than targetOffset - Check if the previous batch contains the target - If there's no previous batch, return the current batch or the previous batch doesn't contain the target, return the current batch 5. After iterating through all batches, check if the last batch contains the target offset. This code path is not thread-safe, so we need to prevent `EOFException`. To avoid this exception, I am still using an early return. In this scenario, `lastOffset` is still used within the loop, but it should be executed at most once within the loop. Therefore, in the new implementation, `lastOffset` will be executed at most once. In most cases, this results in an optimization. Test: Verifying Memory Usage Improvement To evaluate whether this optimization helps, I followed the steps below to monitor memory usage: 1. Start a Standalone Kafka Server ```sh KAFKA_CLUSTER_ID="$(bin/kafka-storage.sh random-uuid)" bin/kafka-storage.sh format --standalone -t $KAFKA_CLUSTER_ID -c config/server.properties bin/kafka-server-start.sh config/server.properties ``` 2. Use Performance Console Tools to Produce and Consume Records Produce Records: ```sh ./kafka-producer-perf-test.sh \ --topic test-topic \ --num-records 1000000000 \ --record-size 100 \ --throughput -1 \ --producer-props bootstrap.servers=localhost:9092 ``` Consume Records: ```sh ./bin/kafka-consumer-perf-test.sh \ --topic test-topic \ --messages 1000000000 \ --bootstrap-server localhost:9092 ``` It can be observed that memory usage has significantly decreased. trunk: ![CleanShot 2025-03-16 at 11 53 31@2x](https://github.com/user-attachments/assets/eec26b1d-38ed-41c8-8c49-e5c68643761b) this PR: ![CleanShot 2025-03-16 at 17 41 56@2x](https://github.com/user-attachments/assets/c8d4c234-18c2-4642-88ae-9f96cf54fccc) Reviewers: Kirk True <kirk@kirktrue.pro>, TengYao Chi <kitingiao@gmail.com>, David Arthur <mumrah@gmail.com>, Jun Rao <junrao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-20 16:33:35 +08:00
Lan Ding	e73719d962	KAFKA-18819 StreamsGroupHeartbeat API and StreamsGroupDescribe API check topic describe (#19183 ) This patch filters out the topic describe unauthorized topics from the StreamsGroupHeartbeat and StreamsGroupDescribe response. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-03-19 20:42:05 +01:00
PoAn Yang	fcca4056fd	KAFKA-18975 Move clients-integration-test out of core module (#19217 ) Move following tests from core to clients-integration-test module. - ClientTelemetryTest - DeleteTopicTest - DescribeAuthorizedOperationsTest - ConsumerIntegrationTest - CustomQuotaCallbackTest - RackAwareAutoTopicCreationTest Move following tests from core to server module. - BootstrapControllersIntegrationTest - LogManagerIntegrationTest Reviewers: Kirk True <kirk@kirktrue.pro>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-20 02:43:19 +08:00
Ritika Reddy	3a3159b01e	KAFKA-18953: [1/N] Add broker side handling for 2 PC (KIP-939) (#19193 ) This patch adds logic to enable and handle two phase commit (2PC) transactions following KIP-939. The changes made are as follows: 1) Add a new broker config called transaction.two.phase.commit.enable which is set to false by default 2) Add new flags enableTwoPCFlag and keepPreparedTxn to handleInitProducerId 3) Return an error if keepPreparedTxn is set to true (for now) Reviewers: Artem Livshits <alivshits@confluent.io>, Justine Olshan <jolshan@confluent.io>	2025-03-19 09:22:00 -07:00
Ken Huang	b805877705	KAFKA-18969 Rewrite ShareConsumerTest#setup and move to clients-integration-tests module (#19202 ) Move share consumer to clients-integration-tests module and use `@BeforeEach` to setup Reviewers: TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-18 14:47:38 +08:00
TengYao Chi	a6a0ea56d8	KAFKA-17171 Add test cases for `STATIC_BROKER_CONFIG`in kraft mode (#18463 ) Given that the `core` module will be separated into other small modules, this test will not be added to the core module. Instead, I added it to the `clients-integration-tests` module since it focuses on the admin client test. The patch should include following test cases. 1. a topic-related static config is added to quorum controller. The configs from topic creation should include it, but `describeConfigs` does not. 2. a topic-related static config is added to quorum controller. The configs from topic creation should include it, and `describeConfigs` does if admin is using controller.bootstrap 3. a topic-related static config is added to broker. The configs from topic creation should NOT include it, but `describeConfigs` does. 4. a topic-related static config is added to broker. The configs from topic creation should NOT include it, and `describeConfigs` does not also if admin is using controller.bootstrap for another, the docs of `STATIC_BROKER_CONFIG` should remind the impact of "controller.properties" BTW, those test cases should leverage new test infra, since new test infra allow us to define configs to broker/controller individually. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-18 00:30:53 +08:00
Bruno Cadonna	a7e40b7c5a	KAFKA-18736: Do not send fields if not needed (#19181 ) The Streams heartbeat request has some fields that are always sent. Those are: - group ID - member ID - member epoch - group instance ID (if static membership is used) Then it has fields that are only sent when joining: - topology and topology epoch - rebalance timeout - process ID - endpoint - client tags Finally, the assignment is only sent if it changed compared to the last sent request. Reviewers: Bill Bejeck <bill@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-16 18:08:56 +01:00
Ken Huang	7bff678699	KAFKA-18859 honor the error message of UnregisterBrokerResponse (#19027 ) Reviewers: Ismael Juma <ismael@juma.me.uk>, TengYao Chi <kitingiao@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-16 03:06:01 +08:00
ClarkChen	e05b0e68e4	KAFKA-18915 Rewrite AdminClientRebootstrapTest to cover the current scenario (#19187 ) Reviewers: Jhen-Yung Hsu <jhenyunghsu@gmail.com>, TengYao Chi <kitingiao@gmail.com>, Ken Huang <s7133700@gmail.com>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-16 02:35:41 +08:00
Kaushik Raina	c32c167e04	KAFKA-18781: Extend RefreshRetriableException related exceptions (#19136 ) - Extended derived exceptions described in [KIP-1050](https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=309496816#KIP1050:ConsistenterrorhandlingforTransactions-RefreshRetriableException) to include the new RefreshRetriableException in base hierarchy. - Added unit tests to validate the hierarchy of the derived exceptions in relevant scenarios. Reviewers: Justine Olshan <jolshan@confluent.io>	2025-03-14 09:11:31 -07:00
Gerard Klijs-Nefkens	b2a01b2754	MINOR: call the serialize method including headers from the MockProducer (#11144 ) Currently when using serializers like the Cloud Event Serializer, we need to do a work around so it doesn't throw an error. Using the method taking the headers would prevent this. Since the default implementation just calls the method without the headers, it's expected to be fully backwards compatible. Reviewers: Divij Vaidya <divijvaidya13@gmail.com>	2025-03-13 18:50:29 +01:00
Mickael Maison	759fbbba8b	KAFKA-14484: Move UnifiedLog to storage module (#19030 ) Rewrite UnifiedLog in Java Reviewers: Jun Rao <jun@confluent.io>, Chia-Ping Tsai <chia7712@gmail.com>	2025-03-13 10:49:55 +01:00
Mickael Maison	55d65cb3ba	MINOR: Cleanups in CoreUtils (#19175 ) Delete unused methods in CoreUtils and switch to Utils.newInstance(). Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-12 19:43:30 +01:00
David Arthur	0ebc3e83c5	MINOR Mar 12 Flaky tests (#19190 ) Mark the following tests as flaky: * StickyAssignorTest > testLargeAssignmentAndGroupWithUniformSubscription * DeleteSegmentsByRetentionTimeTest * QuorumControllerTest > testUncleanShutdownBrokerElrEnabled Reviewers: Andrew Schofield <aschofield@confluent.io>	2025-03-12 13:47:35 -04:00
Abhinav Dixit	c07c59ad24	KAFKA-18932: Removed usage of partition max bytes from share fetch requests (#19148 ) This PR aims to remove the usage of partition max bytes from share fetch requests. Partition Max Bytes is being defined by `PartitionMaxBytesStrategy` which was added to the broker as part of PR https://github.com/apache/kafka/pull/17870 Reviewers: Andrew Schofield <aschofield@confluent.io>, Apoorv Mittal <apoorvmittal10@gmail.com>	2025-03-12 13:19:19 +00:00
David Arthur	701573366f	KAFKA-18933 Add client integration tests module (#19144 ) Adds a new ":clients:integration-test" Gradle module. Relocates one example test from ":core" Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-11 16:36:23 -04:00
David Arthur	903d70d764	MINOR Mark Tls13SelectorTest#testCloseOldestConnection as flaky (#19178 ) This test has a flakiness around 7%. It caused two back-to-back failures on trunk recently. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>	2025-03-11 16:35:38 -04:00
Lucas Brutschy	6551e87815	KAFKA-18925: Add streams groups support to Admin.listGroups (#19155 ) Add support so that Admin.listGroups can represent streams groups and their states. Reviewers: Bill Bejeck <bill@confluent.io>	2025-03-11 15:48:07 +01:00
Bruno Cadonna	59e5890505	KAFKA-18736: Decide when a heartbeat should be sent (#19121 ) This commit adds the conditions to decide when a Streams group heartbeat should be sent. A heartbeat should be sent when: - the group coordinator is available - the member is part of the Streams group or wants to join it - the heartbeat interval expired or the member is leaving the group or acknowledging the assginment This commit does not implement: - not sending fields that did not change - handling errors Reviewers: Zheguang Zhao <zheguang.zhao@alumni.brown.edu>, Lucas Brutschy <lbrutschy@confluent.io>	2025-03-10 17:39:57 +01:00
PoAn Yang	19d8a414ef	KAFKA-15900, KAFKA-18310: fix flaky test testOutdatedCoordinatorAssignment and AbstractCoordinatorTest (#18945 ) Reviewers: Lianet Magrans <lmagrans@confluent.io>	2025-03-10 11:50:35 -04:00
Lucas Brutschy	fc2e3dfce9	MINOR: Disallow unused local variables (#18963 ) Recently, we found a regression that could have been detected by static analysis, since a local variable wasn't being passed to a method during a refactoring, and was left unused. It was fixed in [`7a749b5`](`7a749b589f`), but almost slipped into 4.0. Unused variables are typically detected by IDEs, but this is insufficient to prevent these kinds of bugs. This change enables unused local variable detection in checkstyle for Kafka. A few notes on the usage: - There are two situations in which people actually want to have a local variable but not use it. First, there are `for (Type ignored: collection)` loops which have to loop `collection.length` number of times, but that do not use `ignored` in the loop body. These are typically still easier to read than a classical `for` loop. Second, some IDEs detect it if a return value of a function such as `File.delete` is not being used. In this case, people sometimes store the result in an unused local variable to make ignoring the return value explicit and to avoid the squiggly lines. - In Java 22, unsued local variables can be omitted by using a single underscore `_`. This is supported by checkstyle. In pre-22 versions, IntelliJ allows such variables to be named `ignored` to suppress the unused local variable warning. This pattern is often (but not consistently) used in the Kafka codebase. This is, however, not supported by checkstyle. Since we cannot switch to Java 22, yet, and we want to use automated detection using checkstyle, we have to resort to prefixing the unused local variables with `@SuppressWarnings("UnusedLocalVariable")`. We have to apply this in 11 cases across the Kafka codebase. While not being pretty, I'd argue it's worth it to prevent bugs like the one fixed in [`7a749b5`](`7a749b589f`). Reviewers: Andrew Schofield <aschofield@confluent.io>, David Arthur <mumrah@gmail.com>, Matthias J. Sax <matthias@confluent.io>, Bruno Cadonna <cadonna@apache.org>, Kirk True <ktrue@confluent.io>	2025-03-10 09:37:35 +01:00
Cheryl Simmons	6940bef6e8	MINOR: Small fit and finish changes to Producer config doc strings (#19125 ) - Adding a space, article and punctuation to the Producer config doc strings for consistency and readability. Reviewers: TengYao Chi <kitingiao@gmail.com>, Ken Huang <s7133700@gmail.com>, Justine Olshan <jolshan@confluent.io>	2025-03-07 11:07:35 -08:00
Lucas Brutschy	618ea2c1ca	KAFKA-18285: Add describeStreamsGroup to Admin API (#19116 ) Adds `describeStreamsGroup` to Admin API. This exposes the result of the `DESCRIBE_STREAMS_GROUP` RPC in the Admin API. Reviewers: Bill Bejeck <bill@confluent.io>	2025-03-07 15:56:07 +01:00
David Jacot	8cf2f9a61a	KAFKA-18046; High CPU usage when using Log4j2 (#19138 ) This patch is a first step towards resolving KAFKA-18046. Apache Kafka 4.0 ships with log4j2 so the issue raised in the ticket causing high CPU usage on the fetch path due to LoggerFactory.getLogger() being called on the handling of all fetch responses is not good. Hence, I propose to fix that one by caching the Logger used by the `CompletedFetch` class. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, Ismael Juma <ismael@juma.me.uk>	2025-03-07 00:03:32 -08:00
Matthias J. Sax	d85946da19	MINOR: reduce per-batch logging to TRACE level (#19101 ) Logging on a per-batch bases is very chatty, and should only be done at TRACE level to avoid spamming DEBUG logs. Reviewers: Justine Olshan <jolshan@confluent.io>, Lucas Brutschy <lbrutschy@confluent.io>	2025-03-06 11:06:26 -08:00
Andrew Schofield	1da30bdedf	KAFKA-18900: Experimental share consumer acknowledge mode config (#19113 ) User testing of the `KafkaShareConsumer` interface has revealed some areas which confuse people. One of these is that way that it decides whether you want to use implicit or explicit acknowledgement of records by observing which calls the application issues. We are taking the opportunity to refine the interface before it is finalised. This PR introduces an experimental configuration called `internal.share.acknowledgement.mode` which can be used to make the application declare which kind of acknowledgement it wishes to use. We plan to try out the configuration, assess whether it has helped, and then create a proper consumer configuration that makes this area better. That would require a lot of change in the tests, which explains why this initial PR only has a small number of tests. Reviewers: David Arthur <mumrah@gmail.com>	2025-03-06 17:57:11 +00:00
Ismael Juma	a738df4aaa	KAFKA-18648: Make `records` in `FetchResponse` nullable again (#19131 ) As Jun raised in https://github.com/apache/kafka/pull/18726#discussion_r1972525165, we actually do have a few code paths where `records` remains `null` in the FetchResponse with broker version 3.9 and older: * Compression codec for topic is ZSTD and fetch version < 10: https://github.com/apache/kafka/blob/3.9/core/src/main/scala/kafka/server/KafkaApis.scala#L835 * Down-conversion of zstandard-compressed: https://github.com/apache/kafka/blob/3.9/core/src/main/scala/kafka/server/KafkaApis.scala#L884 * Generic uncaught exception through: https://github.com/apache/kafka/blob/3.9/clients/src/main/java/org/apache/kafka/common/requests/FetchRequest.java#L365 To ensure 4.0 clients don't fail to deserialize fetch responses from brokers with the affected versions, we make `records` nullable again. Reviewers: Chia-Ping Tsai <chia7712@gmail.com>, Jun Rao <junrao@gmail.com>	2025-03-06 09:12:36 -08:00
Alieh Saeedi	7a976c651e	KAFKA-18887: Implement Streams Admin APIs (#19120 ) Implement Admin API extensions beyond list/describe group (delete group, offset-related APIs). * adds methods for describing and manipulating offsets, as described in KIP-1071 * adds corresponding unit tests These are doing the exact same thing as the corresponding consumer group counter-parts. Reviewers: Lucas Brutschy <lbrutschy@confluent.io>	2025-03-06 17:55:21 +01:00

1 2 3 4 5 ...

3902 Commits