The standard advice about student teams is to make them mixed. Mix the abilities, the backgrounds and the disciplines, and the group will teach itself. It appears in almost every teaching handbook and it is right often enough to be a reasonable starting point.
It is also not one decision. A team is mixed or matched on several attributes at once, and the right answer is different for each of them. This article goes attribute by attribute and ends with a set of weights you can defend.

This is the attribute people mean when they say mixed ability, and it is the one with the most contested evidence.
The case for spreading is straightforward. A group where nobody can get started stalls, and the students in it lose the project and the confidence together. Distributing prior knowledge so that every team has at least one person who can make a beginning prevents the worst outcome, which is a group that produces nothing.
The case against is also real. In mixed groups without any structure, work often flows to the most capable member, who does it because it is faster than explaining. That student ends up with more workload and less learning, and the weaker students end up with a grade that reflects someone else's work. The benefit of mixed ability grouping is conditional on the weaker students actually doing something, and that is a design question rather than an allocation question.
What to do. Spread prior knowledge, but pair it with two design features. Structure the project so that tasks are genuinely divisible and assigned, rather than leaving the group to work it out, at least in the early weeks. And track individual contribution through group member evaluation, so that the pattern of one person carrying the group is visible before the end.
A narrower point worth knowing: the range matters more than the mix. A group containing the strongest and the weakest student in the cohort is harder to make work than a group with a moderate spread, because the gap is too large for productive explanation. Balancing algorithms that spread evenly tend to produce moderate ranges, which is one of the quiet advantages of doing this by criteria rather than by hand.
This is the attribute people forget and it causes more group failures than ability ever does.
If four students are free on Tuesday afternoons and the fifth is on placement every Tuesday, the group will either exclude that student or hold meetings nobody can attend. No amount of good intention fixes a timetable clash.
What to do. Cluster hard on availability, campus, mode of study and time zone. In practice this should carry the highest weight in your criteria, higher than skill mix, because it is the constraint that determines whether the group can function at all rather than how well it will function.
Here the useful principle is not heterogeneity for its own sake, it is avoiding the group of one.
A student who is the only woman in an engineering team of five, or the only international student, or the only mature student, participates less, is interrupted more and is more likely to be assigned administrative rather than substantive tasks. When there are two, that effect drops substantially. This is one of the more consistent findings in the research on team composition, and it is the single change in grouping practice most likely to affect equity in your course.
What to do. Set the criterion as a floor rather than as a target. You are not trying to make each group maximally diverse, you are trying to ensure no student is the only one of anything that matters in their team. That is a modest adjustment with an outsized effect, and it is easy to state to students as a design principle.
In interdisciplinary programmes there is a temptation to mix disciplines in every group on principle.
It works when the task genuinely requires more than one specialism, such as a design brief needing both an engineer and a designer. It works poorly when the task is a normal assignment in one field, because the student from the other discipline has nothing to contribute and knows it.
What to do. Mix specialisms only when the assessment brief cannot be completed well by one. If you are mixing disciplines, say explicitly in the brief which part of the work draws on which specialism, because students will not construct that division themselves and the default is that everyone does a bit of everything badly.
There is a persistent belief that groups should balance personality types, usually through some inventory of roles or styles.
The evidence for this is weak. Most of the popular role inventories have poor predictive validity, and allocating on the basis of them adds complexity and survey length for very little return. There is one narrow exception that is worth capturing, which is the difference between students who want to plan everything before starting and students who want to start and adjust. A group split evenly between those two approaches spends its first two weeks arguing about process.
What to do. Skip personality typing. If you want one working style question, ask about planning preference and use it as a low weight criterion, or better, name the difference in the project brief and ask groups to agree a way of working in their first meeting.
A point worth making because it gets less attention than it deserves. Four is the most reliable size for a substantial project. Three works but is fragile, because one dropout leaves a pair and there is not enough slack to absorb a difficult week. Five is workable and is the point at which someone can start to disappear. Six and above and social loafing becomes predictable rather than occasional.
If you are choosing between spending effort on composition and spending it on getting the size right, get the size right first.
A defensible set of criteria for a semester long project might look like this. Availability and practical constraints carry the highest weight, because they are binary. Avoiding isolated members carries the next highest, because the cost is borne by individual students and is invisible in the group's output. Prior relevant experience is spread with a moderate weight. Student preference, if you collect it, carries a low weight so that it influences allocation without dominating it. Everything else is left out.
That takes one survey of six or seven questions and one allocation run, and it produces groups that can meet, can start, and do not leave anyone alone. If you want to know more about how weighting works when criteria pull against each other, our Group Formation FAQ covers it.
Whichever balance you choose, tell students. Two sentences explaining what the groups were optimised for turns the allocation from something arbitrary into something designed, and it gives students a frame for the difficulty they will encounter. A student who knows their group was built to bring together different levels of experience reads the first awkward meeting differently from one who assumes it was random.
If you would like to see how this works inside the tool itself, the Group Formation solution page shows how several weighted criteria are configured and how the groups sync to your LMS. The Group Formation FAQ is a good next stop if you have the practical questions that usually come up before a first run, such as how the activity sits in your LMS, what students see, and what happens when someone misses a deadline.
It is also worth looking at the tools that sit naturally alongside this one, including Group Member Evaluation, Team Based Learning and Peer Review. You can see how they all fit together on the FeedbackFruits tool suite page, or read more about what a feedback and assessment solution actually is if you are building the case for your institution.
And if you want to keep reading on this topic, we have How to form student groups: random, self selected, or criteria based, Create effective surveys for better group selection and Eliminating free riding in group work.